Accuracy is weak on overlapping speech: 33.16% WER with 461 substitutions, 1,976 deletions, and 76 insertions against a 7,579-word reference, ranking 4th of 8 on this input.
What was measured
Output quality
How accurately the returned transcript matches the human reference transcript, measured by WER.
decisive for this rankingtransformation
WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)
What was given, what came back
Test input: Overlapping meeting speech with cross-talk · audio · group: speech-to-text-benchmark
Input — what we sent

33522c2a9c0d4f29b49147f68f855f65.png
0:00 / 0:00
Loading audio...
Overlapping meeting speech with cross-talk
A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.
Why this input is hard
- · overlapping speech
- · background noise robustness
- · multi-speaker separation
- · speaker diarization accuracy
- · long-form audio handling
Output — unretouched


Also checked on this input — same tool, 1 other criterion
Provenance
- Observation
- 67f8ec61-1974-472b-9d73-24058a75af43
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "assemblyai-speech-to-text",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AWS Transcribe⚠ Struggled33.88Transcript quality is weak at 33.88% WER, with 363 substitutions, 2162 deletions, and 43 insertions against 7579 reference words.Deepgram⚠ Struggled36.27On overlapping meeting crosstalk, the transcript quality is weak: WER is 36.27% with 1096 substitutions, 1143 deletions, and 510 insertions against a 7579-word reference, returning 6946 words overall.ElevenLabs Scribe⚠ Struggled26.67Transcript accuracy is weak on crosstalk, with WER 26.67% and 856 substitutions, 782 deletions, and 383 insertions over a 7,579-word reference.Gladia⚠ Struggled37.35On overlapping crosstalk, it still scored the run but WER was 37.35% with 529 substitutions, 2,213 deletions, and 89 insertions against a 7,579-word reference, leaving 5,455 words in the hypothesis.Google Cloud Speech-to-Text⚠ Struggled43.5On overlapping crosstalk, the transcript quality is poor at 43.50% WER, with 633 substitutions, 2614 deletions, and 50 insertions against 7579 reference words, yielding 5015 hypothesis words.GroqCloud◐ MixedProduces no transcript on the oversized file, so WER is unmeasured here; the report explicitly treats accuracy for this input as untested rather than poor.OpenAI Speech-to-Text◐ MixedTranscript accuracy is untested because no transcript was produced; the run stopped at the size rejection before any WER could be measured.Rev AI⚠ Struggled28.33On the overlapping meeting audio, it reached WER 28.33% against a 7579-word reference, with 601 substitutions, 1423 deletions and 123 insertions; it returned 6279 words, so accuracy degraded substantially under crosstalk.Speechmatics⚠ Struggled26.63The transcript is only moderately accurate on overlapping speech, with 26.63% WER (507 substitutions, 1430 deletions, 81 insertions over 7579 reference words) and diarization over-segmenting the meeting into 5 speaker labels.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com