Returns 5703 word-level timed tokens, but no confidence values or speaker labels; the JSON is shallow at depth 3 with 5705 objects.
What was measured
Export
How complete the returned transcript payload is, as reflected in the depth or richness of what the tool outputs.
context, not decisivetransformation
Transcript payload richness is useful for comparison, but it does not measure transcription correctness against the reference. (3 of 3 judges)
What was given, what came back
Test input: Bilingual Spanish-English code-switching speech · audio · group: speech-to-text-benchmark
Input — what we sent
0:00 / 0:00
Loading audio...
Bilingual Spanish-English code-switching speech
A Bangor Miami bilingual corpus recording with mid-sentence switches between Spanish and English. It was used to test multilingual recognition, code-switch detection, and preservation of words across language transitions.
Why this input is hard
- · code-switching detection
- · multilingual language ID
- · mid-sentence language transitions
- · word preservation across language flips
- · hallucination resistance in bilingual speech
Output — unretouched

Also checked on this input — same tool, 1 other criterion
Provenance
- Observation
- 69e1f4f4-c76a-49d2-b48e-787e6b416be6
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "openai-speech-to-text",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Export
AssemblyAI✓ WorkedReturns a rich developer payload with word-level timestamps, confidence values, and speaker labels; the raw response shows 12,181 timed tokens and JSON depth 5 across 12,187 objects, with payload depth 3/3.AWS Transcribe✓ Worked3/3Returns a full developer payload with payload depth 3/3, 6311 word-level timed tokens, confidence values, and speaker labels; the raw JSON walk spans 7 levels and 21039 objects.Deepgram✓ WorkedReturns a rich developer payload on code-switching audio, with word-level timing, confidence, speaker labels, 11,684 timed tokens, 3 distinct speakers, payload depth 3/3, and JSON depth 11 across 11,962 objects.ElevenLabs Scribe✓ Worked3/3The returned payload is rich and fully structured at 3/3 depth, with word-level timestamps, confidence values, speaker labels, and 12,162 timed tokens.Gladia✓ WorkedReturns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 26,166 timed tokens, and 7-level JSON nesting across 26,174 objects.Google Cloud Speech-to-Text◐ MixedReturns a mid-depth transcript payload: payload depth 2/3 with 3688 word-level timed tokens, confidence present, and no speaker labels.GroqCloud✓ WorkedReturns the same verbose JSON shape on code-switching speech, including task, language, duration, segments, word timestamps, confidence data, and no speaker labels; the payload depth is 2/3 with 877 timed tokens.Rev AI✓ WorkedReturned a full developer payload with word-level timing, confidence and speaker labels; the raw response walk reports 5855 timed tokens, payload depth 3/3, and JSON depth 5 across 13384 objects.Speechmatics✓ WorkedReturns a rich transcript payload with payload_depth 3/3, including word-level timing, confidence values, and speaker labels in the JSON response.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com