Transcript quality is strong at 3.63% WER, with 75 substitutions, 12 deletions, and 12 insertions against 2728 reference words.
What was measured
Output quality
How accurately the returned transcript matches the human reference transcript, measured by WER.
decisive for this rankingtransformation
WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)
What was given, what came back
Test input: Medical anatomy narration with dense jargon · audio · group: speech-to-text-benchmark
Input — what we sent
0:00 / 0:00
Loading audio...
Medical anatomy narration with dense jargon
A long narrated excerpt from Henry Gray's Anatomy of the Human Body containing dense medical terminology and accented articulation. It was used to test lexical accuracy on domain-specific jargon and spelling of technical terms.
Why this input is hard
- · domain-specific vocabulary recognition
- · medical term spelling accuracy
- · accented speech robustness
- · phoneme-to-grapheme precision
- · long-form audio handling
Output — unretouched


Also checked on this input — same tool, 1 other criterion
Provenance
- Observation
- 00b55e91-870a-4430-8136-f7b81a548fba
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- observed
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "aws-transcribe",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AssemblyAI✓ WorkedAccuracy is strong on dense medical jargon: 3.78% WER with 71 substitutions, 14 deletions, and 18 insertions against a 2,728-word reference, and the run reports 100.0% jargon recall.Deepgram✓ Worked5.43Keeps lexical accuracy high on dense medical narration: WER is 5.43% with 82 substitutions, 34 deletions, and 32 insertions against a 2728-word reference, and jargon recall is 100.0% (9/9 scored terms).ElevenLabs Scribe✓ Worked3.01Transcript accuracy is strong on dense medical narration, with WER 3.01% and 51 substitutions, 8 deletions, and 23 insertions over a 2,728-word reference.Gladia✓ Worked4.07On dense medical jargon, WER was 4.07% with 73 substitutions, 14 deletions, and 24 insertions against a 2,728-word reference; jargon recall was 100.0% (9/9 scored terms).Google Cloud Speech-to-Text◐ Mixed13.09On dense medical narration, the transcript reaches 13.09% WER with 228 substitutions, 49 deletions, and 80 insertions against 2728 reference words, and it recalls 100.0% of scored jargon terms.GroqCloud✓ Worked3.15Delivers high lexical accuracy on dense medical narration: WER 3.15% with 58 substitutions, 13 deletions, and 15 insertions over 2728 reference words, returning 2730 hypothesis words.OpenAI Speech-to-Text✓ Worked5.17Achieves low transcription error on dense jargon: WER 5.17% with 54 substitutions, 76 deletions, and 11 insertions against 2728 reference words; it returned 2663 words and missed only 'trabeculae'.Rev AI✓ Worked9.79On the dense medical narration, it held WER to 9.79% against a 2728-word reference, with 195 substitutions, 7 deletions and 65 insertions; jargon recall was 77.8% (7/9 scored terms), so it handled the terminology reasonably well but still missed specific terms such as cancellous and trabeculae.Speechmatics✓ Worked3.01The transcript is highly accurate on dense medical narration, with 3.01% WER (49 substitutions, 17 deletions, 16 insertions over 2728 reference words) and 100.0% jargon recall.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com