Speechmatics
Strong batch STT for hard English audio, but weak on code-switching as configured.
Strong batch STT for hard English, but not yet proven multilingual
- You need a batch speech-to-text API that returns structured JSON with word timings, confidence, and speaker labels.
- You are comparing engines on hard English audio such as overlap or jargon-heavy narration.
- You can accept batch transcription rather than live streaming and want competitive per-audio-hour pricing.
- You need proven multilingual or balanced code-switching performance; the configured Spanish-English test was effectively English-heavy and Spanish recall was poor.
Our take
Speechmatics is a strong batch STT default for hard English audio: it won the overlapping-speech test, stayed near the top on medical jargon, and returned structured JSON with word timings, confidence, and speaker labels. The configured bilingual run was English-heavy and showed very weak Spanish recall, so multilingual performance should be re-tested before treating it as proven.
In-Depth Review
Our detailed analysis of Speechmatics — features, performance, and real-world testing.
Feature-by-Feature Breakdown
Asynchronous Batch TranscriptionCompleted all three batch jobs end to end without manual intervention.▾
Feature tested: Asynchronous Batch Transcription
Result: Partial
Verdict: Completed all three batch jobs end to end without manual intervention.
Expected behavior: Speechmatics can transcribe long-form audio as asynchronous batch jobs and return finished transcripts after processing completes. The evidence covers multiple benchmark inputs, including a meeting, a medical narration, and a bilingual conversation.
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk batch run on crosstalk.wav (long-form meeting audio with heavy overlap). — crosstalk.wav
Observed output: Output artifact (Image): Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript. — 07-automation-trace.png
Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk batch run on crosstalk.wav (long-form meeting audio with heavy overlap). — crosstalk.wav
Output artifact: Output artifact (Image): Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript. — 07-automation-trace.png
What changed: Audio file transformed into Image
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-2: Medical Jargon batch run on medical_terms.mp3 (single-speaker anatomical narration). — medical_terms.mp3
Observed output: Output artifact (Image): Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript. — 07-automation-trace-2.png
Input artifact: Input artifact (Audio file): Input-2: Medical Jargon batch run on medical_terms.mp3 (single-speaker anatomical narration). — medical_terms.mp3
Output artifact: Output artifact (Image): Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript. — 07-automation-trace-2.png
What changed: Audio file transformed into Image
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching batch run on mix_language.mp3 (Spanish-English conversation configured with English language routing). — mix_language.mp3
Observed output: Output artifact (Image): Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript. — 07-automation-trace-3.png
Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching batch run on mix_language.mp3 (Spanish-English conversation configured with English language routing). — mix_language.mp3
Output artifact: Output artifact (Image): Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript. — 07-automation-trace-3.png
What changed: Audio file transformed into Image
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Medical Jargon — dense anatomical vocabulary from Gray's Anatomy. — medical_terms.mp3
Observed output: Output artifact (Image): The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them. — 04-transcript-detail-2.png
Input artifact: Input artifact (Audio file): Medical Jargon — dense anatomical vocabulary from Gray's Anatomy. — medical_terms.mp3
Output artifact: Output artifact (Image): The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them. — 04-transcript-detail-2.png
What changed: Audio file transformed into Image
Why it matters / Conclusion: Stable batch execution across all three scored inputs; this benchmark did not measure streaming.
Speechmatics can transcribe long-form audio as asynchronous batch jobs and return finished transcripts after processing completes. The evidence covers multiple benchmark inputs, including a meeting, a medical narration, and a bilingual conversation.




Structured Transcript Metadata OutputConsistently present across all three runs.▾
Feature tested: Structured Transcript Metadata Output
Result: Partial
Verdict: Consistently present across all three runs.
Expected behavior: Speechmatics returns transcript JSON with structured metadata such as word-level timestamps, confidence scores, speaker labels, and language metadata. This was observed across crosstalk, single-speaker medical narration, and bilingual conversation inputs.
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk transcription request. — crosstalk.wav
Observed output: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 7,812 timed tokens in the full run. — 02-response-raw.png
Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk transcription request. — crosstalk.wav
Output artifact: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 7,812 timed tokens in the full run. — 02-response-raw.png
What changed: Audio file transformed into Image
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-2: Medical Jargon transcription request. — medical_terms.mp3
Observed output: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 3,056 timed tokens in the full run. — 02-response-raw-2.png
Input artifact: Input artifact (Audio file): Input-2: Medical Jargon transcription request. — medical_terms.mp3
Output artifact: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 3,056 timed tokens in the full run. — 02-response-raw-2.png
What changed: Audio file transformed into Image
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching transcription request. — mix_language.mp3
Observed output: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 6,907 timed tokens in the full run. — 02-response-raw-3.png
Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching transcription request. — mix_language.mp3
Output artifact: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 6,907 timed tokens in the full run. — 02-response-raw-3.png
What changed: Audio file transformed into Image
Why it matters / Conclusion: Timing metadata was consistently present across all three runs, even when transcription quality varied.
Speechmatics returns transcript JSON with structured metadata such as word-level timestamps, confidence scores, speaker labels, and language metadata. This was observed across crosstalk, single-speaker medical narration, and bilingual conversation inputs.



Speaker DiarizationSpeaker labels are surfaced, but attribution correctness was not scored.▾
Feature tested: Speaker Diarization
Result: Partial
Verdict: Speaker labels are surfaced, but attribution correctness was not scored.
Expected behavior: Speechmatics exposes speaker labels in the transcript output, including on over-segmented crosstalk, single-speaker narration, and mixed-language conversation inputs. The evidence shows label presence and count across runs.
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk diarization test. — crosstalk.wav
Observed output: Output artifact (Image): Transcript detail shows diarization detected 5 labels on a 4-participant crosstalk sample, indicating over-segmentation in the speaker count. — 04-transcript-detail.png
Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk diarization test. — crosstalk.wav
Output artifact: Output artifact (Image): Transcript detail shows diarization detected 5 labels on a 4-participant crosstalk sample, indicating over-segmentation in the speaker count. — 04-transcript-detail.png
What changed: Audio file transformed into Image
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-2: Medical Jargon diarization test. — medical_terms.mp3
Observed output: Output artifact (Image): Raw response preview reports one distinct speaker label on the medical-jargon narration. — 02-response-raw-2.png
Input artifact: Input artifact (Audio file): Input-2: Medical Jargon diarization test. — medical_terms.mp3
Output artifact: Output artifact (Image): Raw response preview reports one distinct speaker label on the medical-jargon narration. — 02-response-raw-2.png
What changed: Audio file transformed into Image
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching diarization test. — mix_language.mp3
Observed output: Output artifact (Image): Raw response preview reports three distinct speaker labels on the bilingual sample. — 02-response-raw-3.png
Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching diarization test. — mix_language.mp3
Output artifact: Output artifact (Image): Raw response preview reports three distinct speaker labels on the bilingual sample. — 02-response-raw-3.png
What changed: Audio file transformed into Image
Why it matters / Conclusion: The tool exposes speaker labeling, but this benchmark does not verify whether those labels were assigned to the correct voices.
Speechmatics exposes speaker labels in the transcript output, including on over-segmented crosstalk, single-speaker narration, and mixed-language conversation inputs. The evidence shows label presence and count across runs.



Multilingual TranscriptionWeak Spanish performance on the bilingual sample as configured.▾
Feature tested: Multilingual Transcription
Result: Partial
Verdict: Weak Spanish performance on the bilingual sample as configured.
Expected behavior: Speechmatics attempts transcription on mixed-language audio, including an English/Spanish sample configured with English as the language setting. The benchmark evidence shows mostly English output with some Spanish tokens missed.
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching test on mix_language.mp3, using English language routing. — mix_language.mp3
Observed output: Output artifact (Image): Error analysis shows the dropped Spanish token 'ahora', 25.06% WER, and only 12.5% Spanish token recall (10 of 80 types) on the bilingual sample. — 04-transcript-detail-3.png
Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching test on mix_language.mp3, using English language routing. — mix_language.mp3
Output artifact: Output artifact (Image): Error analysis shows the dropped Spanish token 'ahora', 25.06% WER, and only 12.5% Spanish token recall (10 of 80 types) on the bilingual sample. — 04-transcript-detail-3.png
What changed: Audio file transformed into Image
Why it matters / Conclusion: Not convincing evidence of balanced multilingual performance; the test was also mostly English, so this needs a fairer retest before strong claims.
Speechmatics attempts transcription on mixed-language audio, including an English/Spanish sample configured with English as the language setting. The benchmark evidence shows mostly English output with some Spanish tokens missed.

How it scored on the research's own criteria
The 4 evaluation dimensions from our hands-on research on Speechmatics, each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.
held up partial failed not exercised by this input
| Criterion | Verdict | What the runs showed | Per input | Proof |
|---|---|---|---|---|
| Output quality | Mixed3/5 | It is excellent on clean medical narration, but the quality drops on harder audio: cross-talk raises error and speaker splitting, while bilingual speech loses most Spanish words. That spread makes the overall transcription quality uneven rather than consistently strong. | open proof ↗ | |
| Automation level | Strong5/5 | Each run finished on its own from submit to transcript fetch, with no manual intervention needed. The only caveat is that the exact step count comes from the documented API flow rather than live call tracing, but the end-to-end automation itself is clearly there. | — | |
| Export | Strong5/5 | It consistently returns a rich transcript package, not just plain text: timed tokens, confidence values, and speaker labels are present on every run. That makes the output easy to feed into downstream tooling without extra reconstruction. | — | |
| Input handling | Strong5/5 | It took every batch we gave it, finished each one cleanly, and reported latency, real-time factor, and list price every time. That is a strong reliability signal rather than a one-off success, even though it runs at the enhanced rate. | open proof ↗ |
Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.
Published pricing
The benchmarked configuration used the Enhanced batch model, which bills above the cheapest headline batch rate.
Vendor pricing was read from Speechmatics' pricing page; benchmark costs in this report are list-price estimates derived from measured duration.
Banner Preview
How the embed badge will look on your site

Embed HTML
Copy this code to your website source
Quick Integration Guide
- 1Copy the HTML code block above.
- 2Paste it into your site's HTML or CMS editor.
- 3Banner appears instantly on your page.
- 4Links back to your tool profile here.
Similar Tools
Discover more AI tools like Speechmatics to enhance your workflow.
Comments (0)
Need a custom AI solution for this use case?
If you are looking to build a custom speech transcription, batch STT, or audio-to-text workflow for your business or internal workflow, email us at contact@futuresmart.ai.
Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.