It was strong on dense medical jargon, but performance dropped on overlapping speech and was middling on bilingual Spanish-English code-switching.
What was measured
Output quality
How accurately the returned transcript matches the human reference transcript, measured by WER.
decisive for this rankingtransformation
WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)
What was given, what came back
Input — what we sent
No input — this is a capability finding
The observation is about the tool itself rather than one test input, so there is nothing to show on this side by design.
Output — unretouched
No output artifact
The verdict rests on the tester's written observation alone — no file was captured for this cell.
Provenance
- Observation
- 9b06ff42-205c-492f-a59c-caceee4ed5ba
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- aggregate-synthesis
- Evidence state
- observed
- Proof shown
- no artifact
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "assemblyai-speech-to-text"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on Output quality
No other tool was measured on this criterion for this input.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com