--- title: "Speechmatics" type: "AI Tool" url: "https://aidemos.com/tools/speechmatics" description: "We fed Speechmatics hard English audio and got structured JSON with word timings, confidence, and speaker labels; code-switching stayed weak." category: "audio-speech" website: "https://docs.speechmatics.com/" published: "2026-08-13T09:18:22.569189+00:00" updated: "2026-09-01T02:56:05.532748+00:00" evidenceCount: 25 verifiedCount: 25 coverage: "dense" --- # Speechmatics Strong batch STT for hard English audio, but weak on code-switching as configured. ## TL;DR Verdict **Mixed but genuinely useful for the right audio** **Where it wins:** - You need batch speech-to-text that returns word timings, confidence values, and speaker labels. - You are comparing engines on hard English audio such as overlapping speakers or jargon-heavy narration. - You can accept batch processing rather than live streaming. **Main limitation:** You need proven balanced multilingual or code-switching performance; this run was mostly English and Spanish recall was poor. **Pricing:** Free $0 — $100 in credit · Pro from $0.129/hr · Enterprise Custom `Batch STT` · `Word timings` · `Speaker labels` · `Weak code-switching` **Website:** [Visit Speechmatics](https://docs.speechmatics.com/) ## Evidence (first-party, tested) *25 tested cells · 25/25 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:speechmatics·cross·automation-level`.* | Criterion | Scenario | Verdict | Score | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | | Automation level | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3d4eff8a9d1b42598661ae09b7b94602.png?v=1) | `ev:speechmatics·cross·automation-level` | | Automation level | Bilingual Spanish-English code-switching speech | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3d4eff8a9d1b42598661ae09b7b94602.png?v=1) | `ev:speechmatics·bilingual-spanish-english-code-switching-speech·automation-level` | | Automation level | Overlapping meeting speech with cross-talk | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/cb24dd68020a401ca46424472a670b33.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-with-cross-talk·automation-level` | | Automation level | Medical anatomy narration with dense jargon | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/cee68b42a3e745f0a8d79c5fb9382c7f.png?v=1) | `ev:speechmatics·medical-anatomy-narration-with-dense-jargon·automation-level` | | Export | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-speechmatics-b1d346bf8bf8.md) | `ev:speechmatics·cross·export` | | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/30ae1948bee84c73bc5674ea3cbc66ee.png?v=1) | `ev:speechmatics·medical-anatomy-narration-with-dense-jargon·export` | | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5da56ed63d0c4c06afbe02abab8daf7b.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-with-cross-talk·export` | | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f3dc6085eddd45adb7fe3500c2d14504.png?v=1) | `ev:speechmatics·bilingual-spanish-english-code-switching-speech·export` | | Export | Overlapping meeting speech and cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a34eedf615cb4bea98956f583bfa9368.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-and-cross-talk·export` | | Export | Medical anatomy dictation with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/dc3b2f4492c74b46b5447129ca663f4f.png?v=1) | `ev:speechmatics·medical-anatomy-dictation-with-dense-jargon·export` | | Export | Spanish-English code-switching conversation | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/c10f76a02b0147d9a217755c97af19c6.png?v=1) | `ev:speechmatics·spanish-english-code-switching-conversation·export` | | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a34eedf615cb4bea98956f583bfa9368.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-with-cross-talk·input-handling` | | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/c10f76a02b0147d9a217755c97af19c6.png?v=1) | `ev:speechmatics·bilingual-spanish-english-code-switching-speech·input-handling` | | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/dc3b2f4492c74b46b5447129ca663f4f.png?v=1) | `ev:speechmatics·medical-anatomy-narration-with-dense-jargon·input-handling` | | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-speechmatics-b1d346bf8bf8.md) | `ev:speechmatics·cross·input-handling` | | Input handling | Overlapping meeting speech and cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a34eedf615cb4bea98956f583bfa9368.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-and-cross-talk·input-handling` | | Input handling | Medical anatomy dictation with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/dc3b2f4492c74b46b5447129ca663f4f.png?v=1) | `ev:speechmatics·medical-anatomy-dictation-with-dense-jargon·input-handling` | | Input handling | Spanish-English code-switching conversation | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/c10f76a02b0147d9a217755c97af19c6.png?v=1) | `ev:speechmatics·spanish-english-code-switching-conversation·input-handling` | | Output quality | Bilingual Spanish-English code-switching speech | ⚠ struggled | 25.1/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f4dec7818bf64a22aa5be35ccc6e7ced.png?v=1) | `ev:speechmatics·bilingual-spanish-english-code-switching-speech·output-quality` | | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 3.0/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/30d646a422644d88b8708f6286b3f190.png?v=1) | `ev:speechmatics·medical-anatomy-narration-with-dense-jargon·output-quality` | | Output quality | Overlapping meeting speech with cross-talk | ◐ mixed | 26.6/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5ecec7585d8741e699fca79c140fce94.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-with-cross-talk·output-quality` | | Output quality | cross-scenario | ◐ mixed | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-speechmatics-b1d346bf8bf8.md) | `ev:speechmatics·cross·output-quality` | | Output quality | Overlapping meeting speech and cross-talk | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a34eedf615cb4bea98956f583bfa9368.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-and-cross-talk·output-quality` | | Output quality | Medical anatomy dictation with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/dc3b2f4492c74b46b5447129ca663f4f.png?v=1) | `ev:speechmatics·medical-anatomy-dictation-with-dense-jargon·output-quality` | | Output quality | Spanish-English code-switching conversation | ⚠ struggled | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/c10f76a02b0147d9a217755c97af19c6.png?v=1) | `ev:speechmatics·spanish-english-code-switching-conversation·output-quality` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. ## Scored across three contrasting audio types Best on overlapping speech, near the top on jargon, and weak on code-switching as configured. - **1/8** Overlapping Speech / Crosstalk — WER 26.63%; best of 8; diarization over-segmented to 5 labels for 4 participants. - **2/10** Medical Jargon — WER 3.01%; second of 10; jargon recall was 100%. - **3/10** Bilingual Code-Switching — WER 25.06%; Spanish token recall was 12.5%; the run was still about 95.5% English. > **Mixed but genuinely useful for the right audio** > > Speechmatics is a strong batch STT choice for hard English audio: it won the overlapping-speech test, stayed near the top on medical jargon, and returned structured JSON with word timings, confidence, and speaker labels. The configured bilingual run was still mostly English and Spanish recall was very low, so multilingual robustness is not proven here. The benchmark also used the Enhanced batch rate, so the quality needs to justify that price against cheaper engines. ## Demo Recording [Video: Speechmatics demo recording](https://cdn.futuresmart.ai/public/aidemos/13b81eaeba7c490f9e9e1ed1b31975c7.mov?v=1) *Video — Screen recording of the STT benchmark folder in Finder, then Terminal running Speechmatics (Ursa/Enhanced) across the three scored inputs and printing benchmark progress and metrics.* ## Feature-by-Feature Breakdown ### Asynchronous Batch Transcription **Verdict:** Reliable batch execution Submits audio as an asynchronous batch job and returns a completed transcript without manual intervention. The member cards exercised crosstalk meeting audio, a jargon-heavy lecture, a medical narration, and a Spanish-English conversation through the documented batch flow. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Stable batch execution across all three inputs. The run completed end to end, but per-call timings were not instrumented, so the exact HTTP call count is documented rather than measured here. ### Transcript Metadata Export **Verdict:** Consistently present across all three runs. Returns transcript payloads with developer-facing metadata such as word-level timestamps, confidence values, speaker labels, transcription_config, language_pack_info, and timed_tokens. The member cards show those structured fields present across the observed runs. **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Timing metadata was consistently present across all three runs, even when transcription quality varied. ### Speaker Diarization **Verdict:** Labels present; attribution unverified Labels speakers in the transcript output. The member cards exercised crosstalk meeting audio and observed speaker tags/labels across runs, including a multi-speaker conversation. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** **Bottom line:** Useful for downstream speaker-aware workflows, but this benchmark did not verify true diarization accuracy. ### Audio Transcription **Verdict:** Weak as configured Transcribes speech audio into a transcript, including hard English narration and Spanish-English mixed-language audio. The member cards cover a hard-English overlap test, a jargon-heavy lecture, and a bilingual code-switching sample. **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Do not treat this run as proof of balanced multilingual performance; it is still mostly English and performs poorly on Spanish tokens. ## Official pricing Pro includes model-tiered batch rates; the benchmarked configuration used Batch Enhanced. | Plan | Price | Notes | | --- | --- | --- | | Free | $0 — $100 in credit | No credit card required; 2 concurrent real-time sessions; 1 batch job/sec. | | Pro ★ (tested) | from $0.129/hr | Batch is tiered by model. The benchmarked configuration used Batch Enhanced at $0.40/hr ($0.006667/min), while Batch Melia 1 is the headline $0.129/hr batch price. | | Enterprise | Custom | No rate limits; custom models; on-prem/container/virtual appliance/on-device; requires sales contact. | *Official pricing page last updated 31 July 2026. The benchmarked run used the Enhanced batch rate, not the cheapest Batch Melia 1 headline rate. The page also lists model-training opt-in and volume discounts; the figures here are list prices.* ## Is It Right For You? **Use it if** - You need batch speech-to-text that returns word timings, confidence values, and speaker labels. - You are comparing engines on hard English audio such as overlapping speakers or jargon-heavy narration. - You can accept batch processing rather than live streaming. **Skip it if** - You need proven balanced multilingual or code-switching performance; this run was mostly English and Spanish recall was poor. - You need measured speaker-attribution accuracy rather than just detected labels and counts. - You need streaming latency results; this benchmark was batch-only. ## Classification - **Category:** audio-speech - **Subcategory:** other-audio-speech - **Type:** text - **Built for:** Other ## Frequently Asked Questions **Q: How accurate was Speechmatics on overlapping speech?** On the crosstalk input, Speechmatics scored WER 26.63% with 507 substitutions, 1430 deletions, and 81 insertions against a 7579-word reference. That was the best score among the engines scored on that input, but it still missed a long overlapping span and over-segmented the meeting into 5 speaker labels for a 4-participant reference. **Q: How did Speechmatics do on medical jargon?** Very well. On the medical-jargon input it scored WER 3.01% with 49 substitutions, 17 deletions, and 16 insertions against a 2728-word reference, and it recalled all scored jargon terms for 100% jargon recall. **Q: How did Speechmatics handle code-switching?** Poorly as configured. The bilingual run had WER 25.06%, but the benchmark notes that the audio was still about 95.5% English overall, and Spanish token recall was only 12.5% (10 of 80 types). **Q: Did Speechmatics return word timestamps and confidence values?** Yes. The raw response and developer-feature summaries show word-level timing, confidence fields, and speaker labels in the payload. **Q: Did the benchmark measure diarization accuracy?** No. It only verified that speaker labels were present and counted them. The report explicitly says attribution correctness was not measured, so diarization accuracy remains unverified. **Q: Was streaming latency measured?** No. This was a batch benchmark. It measured wall-clock latency and real-time factor for batch jobs, but not live streaming latency. **Q: What pricing was used for the benchmarked Speechmatics configuration?** The benchmarked configuration used the Enhanced batch rate at $0.40 per audio-hour ($0.006667/min), not the cheapest Batch Melia 1 headline rate of $0.129/hr. The official pricing page also lists Free and Enterprise tiers, plus model-tiered batch rates. ## Similar Tools AI tools similar to Speechmatics: - [AssemblyAI](https://aidemos.com/tools/assemblyai-speech-to-text) — Fast batch STT with strong metadata and mixed-language performance, but overlap-heavy meetings can drop too many words. - [OpenAI](https://aidemos.com/tools/openai) — Batch speech-to-text with word timestamps, but a strict upload cap and weak multilingual performance make it a mixed fit for hard audio. - [AWS Transcribe](https://aidemos.com/tools/aws-transcribe) — Batch speech-to-text with timestamps and speaker labels, but weak on multilingual audio. - [ElevenLabs Scribe](https://aidemos.com/tools/elevenlabs-scribe) — Fast batch speech-to-text with word-level metadata, strongest on jargon and weaker on overlap/code-switching. - [Rev AI](https://aidemos.com/tools/rev-ai) — Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on hard audio and weak multilingual recall. - [Gladia](https://aidemos.com/tools/gladia) — Batch STT with rich JSON metadata and strong jargon recall, but weak overlap handling and a channel-duplication caveat on bilingual audio. - [Deepgram](https://aidemos.com/tools/deepgram) — Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching. - [ElevenLabs](https://aidemos.com/tools/elevenlabs) — Natural-sounding voice cloning and narration, but with only approximate voice identity. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text transcription, audio indexing, or speech analytics system for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).