Uses a single multipart POST workflow, but this run stops at API rejection and never completes transcription end to end, so the automation is one-step but unsuccessful on this input.
What was measured
Automation level
How many API steps or calls the workflow requires, and whether it completes without operator input.
context, not decisivecapability
How many API steps or whether operator input is needed affects convenience and workflow, but not whether the engine transcribes accurately once run. (3 of 3 judges)
What was given, what came back
Test input: Overlapping meeting speech with cross-talk · audio · group: speech-to-text-benchmark
Input — what we sent
0:00 / 0:00
Loading audio...
Overlapping meeting speech with cross-talk
A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.
Why this input is hard
- · overlapping speech
- · background noise robustness
- · multi-speaker separation
- · speaker diarization accuracy
- · long-form audio handling
Output — unretouched



Also checked on this input — same tool, 2 other criteria
Export✗ FailedReturns no transcript payload at all on the oversized upload; the only returned content is an error object, so export richness is effectively zero.Output quality◐ MixedProduces no transcript on the oversized file, so WER is unmeasured here; the report explicitly treats accuracy for this input as untested rather than poor.
Provenance
- Observation
- e9a2ac7f-48cc-4493-a912-7cc95aadb130
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "groqcloud",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 1 other tool
measured on Automation level
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com