Completes the bilingual run as a single POST with the response returned inline and no operator intervention; the run trace again says per-call timings were not instrumented.
What was measured
Automation level
How many API steps or calls the workflow requires, and whether it completes without operator input.
context, not decisivecapability
How many API steps or whether operator input is needed affects convenience and workflow, but not whether the engine transcribes accurately once run. (3 of 3 judges)
What was given, what came back
Test input: Bilingual Spanish-English code-switching speech · audio · group: speech-to-text-benchmark
Input — what we sent

0:00 / 0:00
Loading audio...
Bilingual Spanish-English code-switching speech
A Bangor Miami bilingual corpus recording with mid-sentence switches between Spanish and English. It was used to test multilingual recognition, code-switch detection, and preservation of words across language transitions.
Why this input is hard
- · code-switching detection
- · multilingual language ID
- · mid-sentence language transitions
- · word preservation across language flips
- · hallucination resistance in bilingual speech
Output — unretouched


Also checked on this input — same tool, 2 other criteria
Export✓ WorkedReturns the same verbose JSON shape on code-switching speech, including task, language, duration, segments, word timestamps, confidence data, and no speaker labels; the payload depth is 2/3 with 877 timed tokens.Output quality⚠ StruggledPerforms much worse on code-switching speech than on medical narration: WER 27.94% with 636 substitutions, 1052 deletions, and 133 insertions over 6517 reference words, plus only 53.8% Spanish token recall.
Provenance
- Observation
- 40fbd740-b1ba-4d5a-a1fd-bfbc6990b8bc
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "groqcloud",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 2 other tools
measured on Automation level
Gladia✓ WorkedThe workflow completes end to end without operator input, using the vendor's documented three-call sequence (upload, submit audio_url to /pre-recorded, then poll result_url until done); this run was multi-stage and did not instrument per-call timings.Google Cloud Speech-to-Text◐ MixedThe recorded workflow submits the audio as a POST to the v2 endpoint and reaches a scored result without operator input, but the trace explicitly says per-call timings were not instrumented, so no measured call count is claimed.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com