Speechmatics icon
audio-speech

Speechmatics

Strong batch STT for hard English audio, but weak on code-switching as configured.

Visit Speechmatics
Batch STTWord-level timingSpeaker labelsHard audio
TL;DR — our verdictUpdated August 2026 · 11 test artifacts

Strong batch STT for hard English, but not yet proven multilingual

Where it wins
  • You need a batch speech-to-text API that returns structured JSON with word timings, confidence, and speaker labels.
  • You are comparing engines on hard English audio such as overlap or jargon-heavy narration.
  • You can accept batch transcription rather than live streaming and want competitive per-audio-hour pricing.
Main limitation
  • You need proven multilingual or balanced code-switching performance; the configured Spanish-English test was effectively English-heavy and Spanish recall was poor.
Pricing (verified plans)
Free $0 — $100 in creditPro (no commitment) from $0.129/hrBatch Melia 1 $0.129/hrBatch Standard $0.24/hr
Strongest test artifacts

Our take

Speechmatics is a strong batch STT default for hard English audio: it won the overlapping-speech test, stayed near the top on medical jargon, and returned structured JSON with word timings, confidence, and speaker labels. The configured bilingual run was English-heavy and showed very weak Spanish recall, so multilingual performance should be re-tested before treating it as proven.

Tutorial recording showing the benchmark folder in Finder and a terminal run selecting Speechmatics [Ursa/Enhanced] and printing WER, latency, RTF, and cost lines.

In-Depth Review

Our detailed analysis of Speechmatics — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Asynchronous Batch Transcription
Completed all three batch jobs end to end without manual intervention.
Test Summary
Feature tested: Asynchronous Batch Transcription
Result: Partial — Completed all three batch jobs end to end without manual intervention.

Feature tested: Asynchronous Batch Transcription

Result: Partial

Verdict: Completed all three batch jobs end to end without manual intervention.

Expected behavior: Speechmatics can transcribe long-form audio as asynchronous batch jobs and return finished transcripts after processing completes. The evidence covers multiple benchmark inputs, including a meeting, a medical narration, and a bilingual conversation.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk batch run on crosstalk.wav (long-form meeting audio with heavy overlap). — crosstalk.wav

Observed output: Output artifact (Image): Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript. — 07-automation-trace.png

Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk batch run on crosstalk.wav (long-form meeting audio with heavy overlap). — crosstalk.wav

Output artifact: Output artifact (Image): Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript. — 07-automation-trace.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-2: Medical Jargon batch run on medical_terms.mp3 (single-speaker anatomical narration). — medical_terms.mp3

Observed output: Output artifact (Image): Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript. — 07-automation-trace-2.png

Input artifact: Input artifact (Audio file): Input-2: Medical Jargon batch run on medical_terms.mp3 (single-speaker anatomical narration). — medical_terms.mp3

Output artifact: Output artifact (Image): Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript. — 07-automation-trace-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching batch run on mix_language.mp3 (Spanish-English conversation configured with English language routing). — mix_language.mp3

Observed output: Output artifact (Image): Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript. — 07-automation-trace-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching batch run on mix_language.mp3 (Spanish-English conversation configured with English language routing). — mix_language.mp3

Output artifact: Output artifact (Image): Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript. — 07-automation-trace-3.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Medical Jargon — dense anatomical vocabulary from Gray's Anatomy. — medical_terms.mp3

Observed output: Output artifact (Image): The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them. — 04-transcript-detail-2.png

Input artifact: Input artifact (Audio file): Medical Jargon — dense anatomical vocabulary from Gray's Anatomy. — medical_terms.mp3

Output artifact: Output artifact (Image): The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them. — 04-transcript-detail-2.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Stable batch execution across all three scored inputs; this benchmark did not measure streaming.

Speechmatics can transcribe long-form audio as asynchronous batch jobs and return finished transcripts after processing completes. The evidence covers multiple benchmark inputs, including a meeting, a medical narration, and a bilingual conversation.

audio
0:00 / 0:00
Loading audio...
Input-1: Overlapping Speech / Crosstalk batch run on crosstalk.wav (long-form meeting audio with heavy overlap).
image
Output artifact for "Asynchronous Batch Transcription" test: Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript., 07-automation-trace.png
Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript.
audio
0:00 / 0:00
Loading audio...
Input-2: Medical Jargon batch run on medical_terms.mp3 (single-speaker anatomical narration).
image
Output artifact for "Asynchronous Batch Transcription" test: Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript., 07-automation-trace-2.png
Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript.
audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching batch run on mix_language.mp3 (Spanish-English conversation configured with English language routing).
image
Output artifact for "Asynchronous Batch Transcription" test: Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript., 07-automation-trace-3.png
Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript.
audio
0:00 / 0:00
Loading audio...
Medical Jargon — dense anatomical vocabulary from Gray's Anatomy.
image
Output artifact for "Asynchronous Batch Transcription" test: The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them., 04-transcript-detail-2.png
The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them.
Bottom Line
Stable batch execution across all three scored inputs; this benchmark did not measure streaming.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Structured Transcript Metadata Output
Consistently present across all three runs.
Test Summary
Feature tested: Structured Transcript Metadata Output
Result: Partial — Consistently present across all three runs.

Feature tested: Structured Transcript Metadata Output

Result: Partial

Verdict: Consistently present across all three runs.

Expected behavior: Speechmatics returns transcript JSON with structured metadata such as word-level timestamps, confidence scores, speaker labels, and language metadata. This was observed across crosstalk, single-speaker medical narration, and bilingual conversation inputs.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk transcription request. — crosstalk.wav

Observed output: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 7,812 timed tokens in the full run. — 02-response-raw.png

Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk transcription request. — crosstalk.wav

Output artifact: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 7,812 timed tokens in the full run. — 02-response-raw.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-2: Medical Jargon transcription request. — medical_terms.mp3

Observed output: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 3,056 timed tokens in the full run. — 02-response-raw-2.png

Input artifact: Input artifact (Audio file): Input-2: Medical Jargon transcription request. — medical_terms.mp3

Output artifact: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 3,056 timed tokens in the full run. — 02-response-raw-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching transcription request. — mix_language.mp3

Observed output: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 6,907 timed tokens in the full run. — 02-response-raw-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching transcription request. — mix_language.mp3

Output artifact: Output artifact (Image): Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 6,907 timed tokens in the full run. — 02-response-raw-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Timing metadata was consistently present across all three runs, even when transcription quality varied.

Speechmatics returns transcript JSON with structured metadata such as word-level timestamps, confidence scores, speaker labels, and language metadata. This was observed across crosstalk, single-speaker medical narration, and bilingual conversation inputs.

audio
0:00 / 0:00
Loading audio...
Input-1: Overlapping Speech / Crosstalk transcription request.
image
Output artifact for "Structured Transcript Metadata Output" test: Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 7,812 timed tokens in the full run., 02-response-raw.png
Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 7,812 timed tokens in the full run.
audio
0:00 / 0:00
Loading audio...
Input-2: Medical Jargon transcription request.
image
Output artifact for "Structured Transcript Metadata Output" test: Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 3,056 timed tokens in the full run., 02-response-raw-2.png
Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 3,056 timed tokens in the full run.
audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching transcription request.
image
Output artifact for "Structured Transcript Metadata Output" test: Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 6,907 timed tokens in the full run., 02-response-raw-3.png
Raw response preview shows the transcription payload with confidence, speaker labels, and word-level timing fields, plus 6,907 timed tokens in the full run.
Bottom Line
Timing metadata was consistently present across all three runs, even when transcription quality varied.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Speaker Diarization
Speaker labels are surfaced, but attribution correctness was not scored.
Test Summary
Feature tested: Speaker Diarization
Result: Partial — Speaker labels are surfaced, but attribution correctness was not scored.

Feature tested: Speaker Diarization

Result: Partial

Verdict: Speaker labels are surfaced, but attribution correctness was not scored.

Expected behavior: Speechmatics exposes speaker labels in the transcript output, including on over-segmented crosstalk, single-speaker narration, and mixed-language conversation inputs. The evidence shows label presence and count across runs.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk diarization test. — crosstalk.wav

Observed output: Output artifact (Image): Transcript detail shows diarization detected 5 labels on a 4-participant crosstalk sample, indicating over-segmentation in the speaker count. — 04-transcript-detail.png

Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk diarization test. — crosstalk.wav

Output artifact: Output artifact (Image): Transcript detail shows diarization detected 5 labels on a 4-participant crosstalk sample, indicating over-segmentation in the speaker count. — 04-transcript-detail.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-2: Medical Jargon diarization test. — medical_terms.mp3

Observed output: Output artifact (Image): Raw response preview reports one distinct speaker label on the medical-jargon narration. — 02-response-raw-2.png

Input artifact: Input artifact (Audio file): Input-2: Medical Jargon diarization test. — medical_terms.mp3

Output artifact: Output artifact (Image): Raw response preview reports one distinct speaker label on the medical-jargon narration. — 02-response-raw-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching diarization test. — mix_language.mp3

Observed output: Output artifact (Image): Raw response preview reports three distinct speaker labels on the bilingual sample. — 02-response-raw-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching diarization test. — mix_language.mp3

Output artifact: Output artifact (Image): Raw response preview reports three distinct speaker labels on the bilingual sample. — 02-response-raw-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: The tool exposes speaker labeling, but this benchmark does not verify whether those labels were assigned to the correct voices.

Speechmatics exposes speaker labels in the transcript output, including on over-segmented crosstalk, single-speaker narration, and mixed-language conversation inputs. The evidence shows label presence and count across runs.

audio
0:00 / 0:00
Loading audio...
Input-1: Overlapping Speech / Crosstalk diarization test.
image
Output artifact for "Speaker Diarization" test: Transcript detail shows diarization detected 5 labels on a 4-participant crosstalk sample, indicating over-segmentation in the speaker count., 04-transcript-detail.png
Transcript detail shows diarization detected 5 labels on a 4-participant crosstalk sample, indicating over-segmentation in the speaker count.
audio
0:00 / 0:00
Loading audio...
Input-2: Medical Jargon diarization test.
image
Output artifact for "Speaker Diarization" test: Raw response preview reports one distinct speaker label on the medical-jargon narration., 02-response-raw-2.png
Raw response preview reports one distinct speaker label on the medical-jargon narration.
audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching diarization test.
image
Output artifact for "Speaker Diarization" test: Raw response preview reports three distinct speaker labels on the bilingual sample., 02-response-raw-3.png
Raw response preview reports three distinct speaker labels on the bilingual sample.
Bottom Line
The tool exposes speaker labeling, but this benchmark does not verify whether those labels were assigned to the correct voices.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark
Multilingual Transcription
Weak Spanish performance on the bilingual sample as configured.
Test Summary
Feature tested: Multilingual Transcription
Result: Partial — Weak Spanish performance on the bilingual sample as configured.

Feature tested: Multilingual Transcription

Result: Partial

Verdict: Weak Spanish performance on the bilingual sample as configured.

Expected behavior: Speechmatics attempts transcription on mixed-language audio, including an English/Spanish sample configured with English as the language setting. The benchmark evidence shows mostly English output with some Spanish tokens missed.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching test on mix_language.mp3, using English language routing. — mix_language.mp3

Observed output: Output artifact (Image): Error analysis shows the dropped Spanish token 'ahora', 25.06% WER, and only 12.5% Spanish token recall (10 of 80 types) on the bilingual sample. — 04-transcript-detail-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching test on mix_language.mp3, using English language routing. — mix_language.mp3

Output artifact: Output artifact (Image): Error analysis shows the dropped Spanish token 'ahora', 25.06% WER, and only 12.5% Spanish token recall (10 of 80 types) on the bilingual sample. — 04-transcript-detail-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Not convincing evidence of balanced multilingual performance; the test was also mostly English, so this needs a fairer retest before strong claims.

Speechmatics attempts transcription on mixed-language audio, including an English/Spanish sample configured with English as the language setting. The benchmark evidence shows mostly English output with some Spanish tokens missed.

audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching test on mix_language.mp3, using English language routing.
image
Output artifact for "Multilingual Transcription" test: Error analysis shows the dropped Spanish token 'ahora', 25.06% WER, and only 12.5% Spanish token recall (10 of 80 types) on the bilingual sample., 04-transcript-detail-3.png
Error analysis shows the dropped Spanish token 'ahora', 25.06% WER, and only 12.5% Spanish token recall (10 of 80 types) on the bilingual sample.
Bottom Line
Not convincing evidence of balanced multilingual performance; the test was also mostly English, so this needs a fairer retest before strong claims.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark

How it scored on the research's own criteria

The 4 evaluation dimensions from our hands-on research on Speechmatics, each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Output qualityMixed3/5It is excellent on clean medical narration, but the quality drops on harder audio: cross-talk raises error and speaker splitting, while bilingual speech loses most Spanish words. That spread makes the overall transcription quality uneven rather than consistently strong.open proof ↗
Automation levelStrong5/5Each run finished on its own from submit to transcript fetch, with no manual intervention needed. The only caveat is that the exact step count comes from the documented API flow rather than live call tracing, but the end-to-end automation itself is clearly there.
ExportStrong5/5It consistently returns a rich transcript package, not just plain text: timed tokens, confidence values, and speaker labels are present on every run. That makes the output easy to feed into downstream tooling without extra reconstruction.
Input handlingStrong5/5It took every batch we gave it, finished each one cleanly, and reported latency, real-time factor, and list price every time. That is a strong reliability signal rather than a one-off success, even though it runs at the enhanced rate.open proof ↗

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Published pricing

The benchmarked configuration used the Enhanced batch model, which bills above the cheapest headline batch rate.

Free
$0 — $100 in credit
No credit card required; 2 concurrent real-time sessions; 1 batch job/sec; service pauses at zero balance until a card is added.
Pro (no commitment)
from $0.129/hr
50 concurrent real-time sessions; 10 batch jobs/sec; billed to the second.
Batch Melia 1
$0.129/hr ($0.00215/min)
Multilingual batch model with mid-conversation language switching; batch only.
Batch Standard
$0.24/hr ($0.004/min)
Strong batch accuracy with lower cost and faster turnaround than Enhanced.
TESTED
Batch Enhanced
$0.40/hr ($0.006667/min)
Highest accuracy across all languages; this is the benchmarked configuration.
Enterprise
Custom
No rate limits; unlimited scale; custom models; on-prem/container/virtual appliance/on-device options.

Vendor pricing was read from Speechmatics' pricing page; benchmark costs in this report are list-price estimates derived from measured duration.

✓ Use This If
You need a batch speech-to-text API that returns structured JSON with word timings, confidence, and speaker labels.
You are comparing engines on hard English audio such as overlap or jargon-heavy narration.
You can accept batch transcription rather than live streaming and want competitive per-audio-hour pricing.
✕ Skip This If
You need proven multilingual or balanced code-switching performance; the configured Spanish-English test was effectively English-heavy and Spanish recall was poor.
You need measured speaker-attribution accuracy rather than just detected speaker labels and counts.
You need streaming latency results; this benchmark was batch-only and did not measure live streaming.
audio-speechaudio-to-textspeechOther
On the overlapping-speech/crosstalk sample, it scored 26.63% WER, which was the best result among the engines scored on that input. It still over-segmented speakers, detecting 5 labels on a 4-participant recording.
It scored 3.01% WER on the medical-jargon narration, which was second best among the scored engines on that input. It also recalled all 9 scored jargon terms.
Poorly in this configuration. The bilingual input was mostly English, and Spanish token recall was only 12.5% (10 of 80 types). The error analysis shows missed Spanish tokens such as "ahora."
Yes. The raw API responses show word-level timestamps, confidence fields, and speaker labels on all three runs.
No. The benchmark detected speaker labels and counts, but it did not score whether each label was assigned to the correct speaker.
No. This was a batch benchmark, so streaming latency was not measured.
The benchmarked configuration used the Enhanced batch model at $0.006667/min, which is $0.40/audio-hour. The vendor also lists cheaper and headline rates such as from $0.129/hr.

Banner Preview

How the embed badge will look on your site

Speechmatics featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/speechmatics?utm_source=speechmatics_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="Speechmatics | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like Speechmatics to enhance your workflow.

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom speech transcription, batch STT, or audio-to-text workflow for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top