Accuracy is middling on code-switching speech: 21.04% WER with 589 substitutions, 652 deletions, and 130 insertions against a 6,517-word reference; Spanish token recall is 72.5% (58/80 types), and the run is best of 10 on this input.
What was measured
Output quality
How accurately the returned transcript matches the human reference transcript, measured by WER.
decisive for this rankingtransformation
WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)
What was given, what came back
Test input: Bilingual Spanish-English code-switching speech · audio · group: speech-to-text-benchmark
Input — what we sent

e319813d52304d89a559933409e837ec.png
0:00 / 0:00
Loading audio...
Bilingual Spanish-English code-switching speech
A Bangor Miami bilingual corpus recording with mid-sentence switches between Spanish and English. It was used to test multilingual recognition, code-switch detection, and preservation of words across language transitions.
Why this input is hard
- · code-switching detection
- · multilingual language ID
- · mid-sentence language transitions
- · word preservation across language flips
- · hallucination resistance in bilingual speech
Output — unretouched


Also checked on this input — same tool, 1 other criterion
Provenance
- Observation
- 9d4f8145-0a65-45a9-bc6d-0299eef9e293
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "assemblyai-speech-to-text",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AWS Transcribe◐ Mixed23.06Transcript quality is middling at 23.06% WER, with 673 substitutions, 665 deletions, and 165 insertions against 6517 reference words; Spanish recall is only 30.0% (24/80 types).Deepgram✗ Failed38.13Breaks down on code-switching speech: WER is 38.13% with 793 substitutions, 1259 deletions, and 433 insertions against a 6517-word reference, and Spanish token recall is only 3.8% (3/80 types).ElevenLabs Scribe⚠ Struggled28.57Transcript accuracy is weak on code-switching speech, with WER 28.57% and 752 substitutions, 778 deletions, and 332 insertions over a 6,517-word reference.Gladia✗ Failed88.45On bilingual code-switching, the response duplicated across 2 channels, inflating output to 10,765 words against a 6,517-word reference and making the reported 88.45% WER an artefact of duplication rather than a usable accuracy score.Google Cloud Speech-to-Text✗ Failed56.13On code-switching speech, the transcript quality is very poor at 56.13% WER, with 716 substitutions, 2884 deletions, and 58 insertions against 6517 reference words, and only 5.0% Spanish token recall (4/80 types).GroqCloud⚠ Struggled27.94Performs much worse on code-switching speech than on medical narration: WER 27.94% with 636 substitutions, 1052 deletions, and 133 insertions over 6517 reference words, plus only 53.8% Spanish token recall.OpenAI Speech-to-Text⚠ Struggled26.12Transcription quality degrades sharply on code-switching: WER 26.12% with 623 substitutions, 947 deletions, and 132 insertions against 6517 reference words; Spanish token recall is only 46.2% (37/80).Rev AI⚠ Struggled25.16On the code-switching sample, it reached WER 25.16% against a 6517-word reference, with 691 substitutions, 793 deletions and 156 insertions; Spanish token recall was only 8.8% (7/80 types), so it dropped or anglicised most Spanish material.Speechmatics⚠ Struggled25.06The transcript struggles on code-switching speech, with 25.06% WER (560 substitutions, 981 deletions, 92 insertions over 6517 reference words) and only 12.5% Spanish token recall (10/80 types).
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com