Completes transcription as a single POST with the response returned inline and no operator intervention; the run trace notes that per-call timings were not instrumented.
What was measured
Automation level
How many API steps or calls the workflow requires, and whether it completes without operator input.
context, not decisivecapability
How many API steps or whether operator input is needed affects convenience and workflow, but not whether the engine transcribes accurately once run. (3 of 3 judges)
What was given, what came back
Test input: Medical anatomy narration with dense jargon · audio · group: speech-to-text-benchmark
Input — what we sent

0:00 / 0:00
Loading audio...
Medical anatomy narration with dense jargon
A long narrated excerpt from Henry Gray's Anatomy of the Human Body containing dense medical terminology and accented articulation. It was used to test lexical accuracy on domain-specific jargon and spelling of technical terms.
Why this input is hard
- · domain-specific vocabulary recognition
- · medical term spelling accuracy
- · accented speech robustness
- · phoneme-to-grapheme precision
- · long-form audio handling
Output — unretouched


Also checked on this input — same tool, 2 other criteria
Export✓ WorkedReturns a verbose JSON transcript payload with task, language, duration, segments, word timestamps, confidence data, and no speaker labels; the payload depth is 2/3 with 206 timed tokens.Output quality✓ WorkedDelivers high lexical accuracy on dense medical narration: WER 3.15% with 58 substitutions, 13 deletions, and 15 insertions over 2728 reference words, returning 2730 hypothesis words.
Provenance
- Observation
- 7ca3e91d-1ecd-4090-b774-95b3fdcace36
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "groqcloud",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 1 other tool
measured on Automation level
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com