Returns a mid-depth transcript payload: payload depth 2/3 with 2750 word-level timed tokens, confidence present, and no speaker labels.
What was measured
Export
How complete the returned transcript payload is, as reflected in the depth or richness of what the tool outputs.
context, not decisivetransformation
Transcript payload richness is useful for comparison, but it does not measure transcription correctness against the reference. (3 of 3 judges)
What was given, what came back
Test input: Medical anatomy narration with dense jargon · audio · group: speech-to-text-benchmark
Input — what we sent
0:00 / 0:00
Loading audio...
Medical anatomy narration with dense jargon
A long narrated excerpt from Henry Gray's Anatomy of the Human Body containing dense medical terminology and accented articulation. It was used to test lexical accuracy on domain-specific jargon and spelling of technical terms.
Why this input is hard
- · domain-specific vocabulary recognition
- · medical term spelling accuracy
- · accented speech robustness
- · phoneme-to-grapheme precision
- · long-form audio handling
Also checked on this input — same tool, 2 other criteria
Automation level◐ MixedThe recorded workflow submits the audio as a POST to the v2 endpoint and reaches a scored result without operator input, but the trace explicitly says per-call timings were not instrumented, so no measured call count is claimed.Output quality◐ MixedOn dense medical narration, the transcript reaches 13.09% WER with 228 substitutions, 49 deletions, and 80 insertions against 2728 reference words, and it recalls 100.0% of scored jargon terms.
Provenance
- Observation
- edb2573e-59c0-4979-a4fe-8f0f627f3acd
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "google-cloud-speech-to-text",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Export
AssemblyAI✓ WorkedReturns a rich developer payload with word-level timestamps, confidence values, and speaker labels; the raw response shows 5,437 timed tokens and JSON depth 5 across 5,443 objects, with payload depth 3/3.AWS Transcribe✓ Worked3/3Returns a full developer payload with payload depth 3/3, 2829 word-level timed tokens, confidence values, and speaker labels; the raw JSON walk spans 7 levels and 9023 objects.Deepgram✓ WorkedReturns a rich developer payload on medical narration, with word-level timing, confidence, speaker labels, 5,845 timed tokens, 1 distinct speaker, payload depth 3/3, and JSON depth 11 across 5,881 objects.ElevenLabs Scribe✓ Worked3/3The returned payload is rich and fully structured at 3/3 depth, with word-level timestamps, confidence values, speaker labels, and 5,448 timed tokens.Gladia✓ WorkedReturns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 5,854 timed tokens, and 7-level JSON nesting across 5,862 objects.GroqCloud✓ WorkedReturns a verbose JSON transcript payload with task, language, duration, segments, word timestamps, confidence data, and no speaker labels; the payload depth is 2/3 with 206 timed tokens.OpenAI Speech-to-Text◐ MixedReturns a fairly rich verbose_json payload with 2664 word-level timed tokens, but no confidence values and no speaker labels; JSON depth is 3 with 2666 objects.Rev AI✓ WorkedReturned a full developer payload with word-level timing, confidence and speaker labels; the raw response walk reports 2779 timed tokens, payload depth 3/3, and JSON depth 5 across 5873 objects.Speechmatics✓ WorkedReturns a rich transcript payload with payload_depth 3/3, including word-level timing, confidence values, and speaker labels in the JSON response.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
