--- title: "Speechmatics" type: "AI Tool" url: "https://aidemos.com/tools/speechmatics" description: "In our hard-English STT tests, Speechmatics won overlapping speech and returned structured JSON, but bilingual Spanish recall was very weak." category: "audio-speech" website: "https://docs.speechmatics.com/" published: "2026-08-13T09:18:22.569189+00:00" updated: "2026-08-20T12:45:51.695472+00:00" evidenceCount: 12 verifiedCount: 12 coverage: "dense" --- # Speechmatics Strong batch STT for hard English audio, but weak on code-switching as configured. ## TL;DR Verdict **Strong batch STT for hard English, but not yet proven multilingual** **Where it wins:** - You need a batch speech-to-text API that returns structured JSON with word timings, confidence, and speaker labels. - You are comparing engines on hard English audio such as overlap or jargon-heavy narration. - You can accept batch transcription rather than live streaming and want competitive per-audio-hour pricing. **Main limitation:** You need proven multilingual or balanced code-switching performance; the configured Spanish-English test was effectively English-heavy and Spanish recall was poor. **Pricing:** Free $0 — $100 in credit · Pro (no commitment) from $0.129/hr · Batch Melia 1 $0.129/hr ($0.00215/min) · Batch Standard $0.24/hr ($0.004/min) `Batch STT` · `Word-level timing` · `Speaker labels` · `Hard audio` **Website:** [Visit Speechmatics](https://docs.speechmatics.com/) ## Evidence (first-party, tested) *12 tested cells · 12/12 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:speechmatics·cross·automation-level`.* | Criterion | Scenario | Verdict | Score | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | | Automation level | cross-scenario | ◐ mixed | 3/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/cb24dd68020a401ca46424472a670b33.png?v=1) | `ev:speechmatics·cross·automation-level` | | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/288f89662ee148859ab56bf4dab43819.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-with-cross-talk·export` | | Export | cross-scenario | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/2b937b7cbe3f40ad9a5bcfe854ee64b5.jpeg?v=1) | `ev:speechmatics·cross·export` | | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/116616caa7934514adc93166c40ef809.png?v=1) | `ev:speechmatics·bilingual-spanish-english-code-switching-speech·export` | | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/30ae1948bee84c73bc5674ea3cbc66ee.png?v=1) | `ev:speechmatics·medical-anatomy-narration-with-dense-jargon·export` | | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/dc3b2f4492c74b46b5447129ca663f4f.png?v=1) | `ev:speechmatics·medical-anatomy-narration-with-dense-jargon·input-handling` | | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/8e905f5ec1cb456e83f24f54dc7243c5.png?v=1) | `ev:speechmatics·bilingual-spanish-english-code-switching-speech·input-handling` | | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/fa4ab6670e9a407fb10bd06f20c32128.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-with-cross-talk·input-handling` | | Input handling | cross-scenario | ◐ mixed | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-speechmatics-b1d346bf8bf8.md) | `ev:speechmatics·cross·input-handling` | | Output quality | Overlapping meeting speech with cross-talk | ◐ mixed | 26.6/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/c0bbec2f8f474e738653322148c03d62.png?v=1) | `ev:speechmatics·overlapping-meeting-speech-with-cross-talk·output-quality` | | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 3.0/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/69ec95086bfb4766871b108d84895f76.png?v=1) | `ev:speechmatics·medical-anatomy-narration-with-dense-jargon·output-quality` | | Output quality | Bilingual Spanish-English code-switching speech | ⚠ struggled | 25.1/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/bebd78737db04320b7db6f8b759418c6.png?v=1) | `ev:speechmatics·bilingual-spanish-english-code-switching-speech·output-quality` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. > **Strong batch STT for hard English, but not yet proven multilingual** > > Speechmatics is a strong batch STT default for hard English audio: it won the overlapping-speech test, stayed near the top on medical jargon, and returned structured JSON with word timings, confidence, and speaker labels. The configured bilingual run was English-heavy and showed very weak Spanish recall, so multilingual performance should be re-tested before treating it as proven. ## Demo Recording [Video: Speechmatics demo recording](https://cdn.futuresmart.ai/public/aidemos/8d255843e4524c79a76bb02f63a37e7b.mov?v=1) *Video — Tutorial recording showing the benchmark folder in Finder and a terminal run selecting Speechmatics [Ursa/Enhanced] and printing WER, latency, RTF, and cost lines.* ## Feature-by-Feature Breakdown ### Asynchronous Batch Transcription **Verdict:** Completed all three batch jobs end to end without manual intervention. Speechmatics can transcribe long-form audio as asynchronous batch jobs and return finished transcripts after processing completes. The evidence covers multiple benchmark inputs, including a meeting, a medical narration, and a bilingual conversation. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Stable batch execution across all three scored inputs; this benchmark did not measure streaming. ### Structured Transcript Metadata Output **Verdict:** Consistently present across all three runs. Speechmatics returns transcript JSON with structured metadata such as word-level timestamps, confidence scores, speaker labels, and language metadata. This was observed across crosstalk, single-speaker medical narration, and bilingual conversation inputs. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Timing metadata was consistently present across all three runs, even when transcription quality varied. ### Speaker Diarization **Verdict:** Speaker labels are surfaced, but attribution correctness was not scored. Speechmatics exposes speaker labels in the transcript output, including on over-segmented crosstalk, single-speaker narration, and mixed-language conversation inputs. The evidence shows label presence and count across runs. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** The tool exposes speaker labeling, but this benchmark does not verify whether those labels were assigned to the correct voices. ### Multilingual Transcription **Verdict:** Weak Spanish performance on the bilingual sample as configured. Speechmatics attempts transcription on mixed-language audio, including an English/Spanish sample configured with English as the language setting. The benchmark evidence shows mostly English output with some Spanish tokens missed. **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Not convincing evidence of balanced multilingual performance; the test was also mostly English, so this needs a fairer retest before strong claims. ## Published pricing The benchmarked configuration used the Enhanced batch model, which bills above the cheapest headline batch rate. | Plan | Price | Notes | | --- | --- | --- | | Free | $0 — $100 in credit | No credit card required; 2 concurrent real-time sessions; 1 batch job/sec; service pauses at zero balance until a card is added. | | Pro (no commitment) | from $0.129/hr | 50 concurrent real-time sessions; 10 batch jobs/sec; billed to the second. | | Batch Melia 1 | $0.129/hr ($0.00215/min) | Multilingual batch model with mid-conversation language switching; batch only. | | Batch Standard | $0.24/hr ($0.004/min) | Strong batch accuracy with lower cost and faster turnaround than Enhanced. | | Batch Enhanced ★ (tested) | $0.40/hr ($0.006667/min) | Highest accuracy across all languages; this is the benchmarked configuration. | | Enterprise | Custom | No rate limits; unlimited scale; custom models; on-prem/container/virtual appliance/on-device options. | *Vendor pricing was read from Speechmatics' pricing page; benchmark costs in this report are list-price estimates derived from measured duration.* ## Is It Right For You? **Use it if** - You need a batch speech-to-text API that returns structured JSON with word timings, confidence, and speaker labels. - You are comparing engines on hard English audio such as overlap or jargon-heavy narration. - You can accept batch transcription rather than live streaming and want competitive per-audio-hour pricing. **Skip it if** - You need proven multilingual or balanced code-switching performance; the configured Spanish-English test was effectively English-heavy and Spanish recall was poor. - You need measured speaker-attribution accuracy rather than just detected speaker labels and counts. - You need streaming latency results; this benchmark was batch-only and did not measure live streaming. ## Classification - **Category:** audio-speech - **Subcategory:** audio-to-text - **Type:** speech - **Built for:** Other ## Frequently Asked Questions **Q: How accurate was Speechmatics on overlapping speech?** On the overlapping-speech/crosstalk sample, it scored 26.63% WER, which was the best result among the engines scored on that input. It still over-segmented speakers, detecting 5 labels on a 4-participant recording. **Q: How did Speechmatics do on medical jargon?** It scored 3.01% WER on the medical-jargon narration, which was second best among the scored engines on that input. It also recalled all 9 scored jargon terms. **Q: How did Speechmatics handle code-switching?** Poorly in this configuration. The bilingual input was mostly English, and Spanish token recall was only 12.5% (10 of 80 types). The error analysis shows missed Spanish tokens such as "ahora." **Q: Did Speechmatics return word timestamps and confidence values?** Yes. The raw API responses show word-level timestamps, confidence fields, and speaker labels on all three runs. **Q: Did the benchmark measure diarization accuracy?** No. The benchmark detected speaker labels and counts, but it did not score whether each label was assigned to the correct speaker. **Q: Was streaming latency measured?** No. This was a batch benchmark, so streaming latency was not measured. **Q: What pricing was used for the benchmarked Speechmatics configuration?** The benchmarked configuration used the Enhanced batch model at $0.006667/min, which is $0.40/audio-hour. The vendor also lists cheaper and headline rates such as from $0.129/hr. ## Similar Tools AI tools similar to Speechmatics: - [Deepgram Nova-3](https://aidemos.com/tools/deepgram-nova-3) — Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching. - [Rev AI](https://aidemos.com/tools/rev-ai) — Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on hard audio and weak multilingual recall. - [Gladia](https://aidemos.com/tools/gladia) — Batch STT with rich JSON metadata and strong jargon recall, but weak overlap handling and a channel-duplication caveat on bilingual audio. - [ElevenLabs](https://aidemos.com/tools/elevenlabs) — Natural-sounding voice cloning and narration, but with only approximate voice identity. ## Need a custom AI solution for this use case? If you are looking to build a custom speech transcription, batch STT, or audio-to-text workflow for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).