--- title: "Best AI Tools for Accurate Speech-to-Text on Hard Audio" type: "Ranking" url: "https://aidemos.com/best/speech-to-text-apis" description: "Developers choosing a speech-to-text engine need more than clean-audio demos: they need to know which system holds up on overlapping speakers, technical terms, and code-switching, while still returning timestamps, speaker labels, low latency, and sensible cost. We benchmarked 10 engines on the same three long real-world recordings and compared WER, diarization, timestamp payload depth, runtime, and price." readTime: "13 min read" tested: "AssemblyAI (Universal) vs Speechmatics (Ursa/Enhanced) vs ElevenLabs Scribe vs AWS Transcribe vs Deepgram Nova-3 vs Gladia vs Rev AI vs GroqCloud (Whisper Large-v3) vs OpenAI vs Google Cloud STT v2" testedDate: "2026-08-23" category: "audio-speech" published: "2026-08-13T09:18:22.578829+00:00" updated: "2026-09-02T19:13:04.274934+00:00" evidenceCount: 72 verifiedCount: 61 coverage: "dense" --- # Best AI Tools for Accurate Speech-to-Text on Hard Audio `10 Tools Tested` · `3 Hard Inputs` · `WER Benchmark` · `Speaker Diarization` · `Word Timestamps` · `Long-Form Audio` · `Code-Switching` **Tested:** AssemblyAI (Universal) vs Speechmatics (Ursa/Enhanced) vs ElevenLabs Scribe vs AWS Transcribe vs Deepgram Nova-3 vs Gladia vs Rev AI vs GroqCloud (Whisper Large-v3) vs OpenAI vs Google Cloud STT v2 · 2026-08-23 > Developers choosing a speech-to-text engine need more than clean-audio demos: they need to know which system holds up on overlapping speakers, technical terms, and code-switching, while still returning timestamps, speaker labels, low latency, and sensible cost. We benchmarked 10 engines on the same three long real-world recordings and compared WER, diarization, timestamp payload depth, runtime, and price. ## Our Verdict **#1 pick: AssemblyAI (Universal)** (Best) — The strongest overall pick if you need a long-form transcription engine that stays competitive on the hardest inputs and still returns speaker labels and word timestamps. - #2 Speechmatics (Ursa/Enhanced) — Strong payload export and clean batch handling, but multilingual and crosstalk accuracy are the weak spots. - #3 ElevenLabs Scribe — Reliable single-call STT with rich output; strongest on dense medical narration, weaker on crosstalk and code-switching. - #4 AWS Transcribe — Reliable batch transcription with rich payloads, but uneven accuracy on difficult speech. - #5 Deepgram Nova-3 — Strong full-payload batch STT, but shaky on overlap-heavy and code-switching speech. - #6 Gladia — Strong batch workflow with rich transcript payloads, but it loses ground on hard speech and channel duplication. - #7 Rev AI — Cheap, fully automated batch transcription with rich payloads, but accuracy drops sharply on crosstalk and code-switching. - #8 GroqCloud (Whisper Large-v3) — Strong on dense narration, but bounded by a 25 MB upload cap and weaker on code-switching. - #9 OpenAI — Good on clean narration, but it falls apart on code-switching and refuses files over the upload cap. - #10 Google Cloud STT v2 — Reliable at taking long batch audio through to a result, but transcript accuracy swings sharply by audio type. ## How We Tested We ran the same three long-form audio clips through every available transcription engine using the benchmark harness, then scored each transcript against a human-verified reference with WER. The suite was built to isolate three hard modes of STT: overlapping speakers and room noise, dense medical jargon, and Spanish-English code-switching. For each run we recorded wall-clock latency, real-time factor, cost, and whether the engine returned word timestamps, confidence, and speaker labels. This is a batch/asynchronous benchmark only; it does not include a clean baseline or a direct streaming comparison run. **What we evaluated:** | Criterion | Description | | --- | --- | | Automation level | How many API steps or calls the workflow requires, and whether it completes without operator input. | | Export | How complete the returned transcript payload is, as reflected in the depth or richness of what the tool outputs. | | Input handling | Whether the tool accepts the benchmark audio as provided and processes it end to end without objection, including the file size/duration it can handle. | | Output quality | How accurately the returned transcript matches the human reference transcript, measured by WER. | ## The Ranking 10 tools tested head-to-head on the same input. ### 1. AssemblyAI (Universal) — Best *Strong batch API with rich transcript payloads, but uneven accuracy on overlapping speech.* The strongest overall pick if you need a long-form transcription engine that stays competitive on the hardest inputs and still returns speaker labels and word timestamps. ### 2. Speechmatics (Ursa/Enhanced) — Usable *Strong payload export and clean batch handling, but multilingual and crosstalk accuracy are the weak spots.* A top-tier choice when crosstalk is the hardest problem and you still want full developer payloads, but its bilingual performance is less convincing than AssemblyAI's. ### 3. ElevenLabs Scribe — Usable *Reliable single-call STT with rich output; strongest on dense medical narration, weaker on crosstalk and code-switching.* Excellent speed and strong results on the medical clip, with complete timestamps and speaker labels, but speaker over-segmentation and weaker code-switching keep it below the top two. ### 4. AWS Transcribe — Usable *Reliable batch transcription with rich payloads, but uneven accuracy on difficult speech.* A solid, broadly usable batch engine with good diarization and timestamps, though it is slower than the leaders and not as strong on the hardest multilingual clip. ### 5. Deepgram Nova-3 — Usable *Strong full-payload batch STT, but shaky on overlap-heavy and code-switching speech.* Useful when speed matters, especially on clean jargon audio, but the overlap and bilingual results are too uneven to call it a safe default for hard speech. ### 6. Gladia — Usable *Strong batch workflow with rich transcript payloads, but it loses ground on hard speech and channel duplication.* A very fast and accurate medical-jargon engine, but the bilingual run duplicated channels and wrecked the transcript, so it needs a channel-handling fix before it is dependable. ### 7. Rev AI — Usable *Cheap, fully automated batch transcription with rich payloads, but accuracy drops sharply on crosstalk and code-switching.* A workable async option that handles overlap and multilingual speech reasonably well, but its medical-jargon performance and speaker count drift keep it behind the front-runners. ### 8. GroqCloud (Whisper Large-v3) — Usable *Strong on dense narration, but bounded by a 25 MB upload cap and weaker on code-switching.* Promising on the files it accepts, yet the 25 MB cap rejected the longest meeting clip, which makes it risky for the long-form benchmark this page is about. ### 9. OpenAI — Usable *Good on clean narration, but it falls apart on code-switching and refuses files over the upload cap.* Good transcription on the shorter files, but the long overlap clip was rejected and the developer payload is thinner than the leaders. ### 10. Google Cloud STT v2 — Usable *Reliable at taking long batch audio through to a result, but transcript accuracy swings sharply by audio type.* It returned transcripts, but the combination of poor WER, no usable diarization here, and very high latency leaves it well behind the best options. ## Full Breakdown ### AssemblyAI (Universal) Most balanced overall: strong on overlap, strong on jargon, and competitive on the bilingual clip, with full word timestamps and speaker labels. ![AssemblyAI (Universal) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![AssemblyAI (Universal) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![AssemblyAI (Universal) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It handled all three inputs, returned the full developer payload on each run, and kept bilingual accuracy much better than the weakest engines. The overlap run still had substantial deletion error, but the engine stayed usable and the speaker count matched the underlying conversation structure on the hardest meeting clip. **Where it struggled:** - The overlap clip still produced a large missing span, and the bilingual clip lost a noticeable amount of Spanish despite a strong 72.5% token recall. It is good overall, but not flawless on the hardest acoustic and multilingual edges. **What came out:** ![AssemblyAI (Universal) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/b0f36d9c22e145e0b521882a7d8c127e.png?v=1) *Output — The pipeline sent a POST request to the AssemblyAI endpoint with speaker labels, punctuation, and text formatting enabled for crosstalk.wav.* ![AssemblyAI (Universal) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/6475a5e628f6469bb90133294fa31ca6.png?v=1) *Output — The crosstalk run scored WER 33.16%, with 461 substitutions, 1976 deletions, 76 insertions, 35.22s latency, and $0.13688 cost.* ![AssemblyAI (Universal) output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/d026cc838f184cafa97048efceea118d.png?v=1) *Output — The largest crosstalk divergence showed a long missing span in the overlap-heavy section, even though the engine still returned 4 speaker labels for 4 participants.* ![AssemblyAI (Universal) output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/d987485f2bf34280b1c98566d97f44ab.png?v=1) *Output — The raw response began successfully and included word-level timing, confidence, and speaker labels; 11743 timed tokens were detected in the payload.* ![AssemblyAI (Universal) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/7b2e0ac4924444d39c8bd690f2135982.png?v=1) *Output — The crosstalk job completed end to end as a multi-stage async flow, returning 5679 words from a 7579-word reference.* ![AssemblyAI (Universal) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/881ffe23f38d4c67838ee72daf7f9e2c.png?v=1) *Output — The medical-jargon run submitted medical_terms.mp3 with speaker labels, punctuation, and formatting enabled.* ![AssemblyAI (Universal) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/1fcfaa34ccc342b1bfba913ed7d640c5.png?v=1) *Output — The medical-jargon run scored WER 3.78%, with perfect jargon recall and 20.14s latency.* ![AssemblyAI (Universal) output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/c0c8e62994434513bd45f97516766135.png?v=1) *Output — The main medical-term error was a small wording difference around 'all together' versus 'altogether', not a vocabulary failure.* ![AssemblyAI (Universal) output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/35765f81b42e4449bfa155a3ddae65d3.png?v=1) *Output — The response showed successful transcription of the medical clip with word-level timestamps and one speaker label.* ![AssemblyAI (Universal) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/366ff854b6264f0e9c29e6751d9b9f30.png?v=1) *Output — The medical-jargon job completed cleanly and returned 2732 words against a 2728-word reference.* ![AssemblyAI (Universal) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/3e05f4976b8a4626b5c94d85bbcd5c9d.png?v=1) *Output — The bilingual run submitted mix_language.mp3 with diarization enabled and a two-channel audio payload.* ![AssemblyAI (Universal) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/54fefc4470df44c393815aa2eb2f0970.png?v=1) *Output — The bilingual run scored WER 21.04%, with 589 substitutions, 652 deletions, 130 insertions, and 72.5% Spanish token recall.* ![AssemblyAI (Universal) output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/76864ef504d34f068d7a4373959e823c.png?v=1) *Output — The bilingual error site dropped the Spanish phrase around 'mi entoces ahora', but the engine still retained the overall sentence structure better than most peers.* ![AssemblyAI (Universal) output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/074b7aeb60c84b3f9b6ff39ef92961a3.png?v=1) *Output — The raw response showed word-level timing, confidence, speaker labels, and 12181 timed tokens for the bilingual clip.* ![AssemblyAI (Universal) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/985619ffd59441038ff310bea3517c28.png?v=1) *Output — The bilingual job completed end to end and returned 5995 words from a 6517-word reference.* ### Speechmatics (Ursa/Enhanced) Best on overlap and very strong on medical jargon, with full timing/confidence/speaker output and sensible batch pricing. ![Speechmatics (Ursa/Enhanced) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![Speechmatics (Ursa/Enhanced) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![Speechmatics (Ursa/Enhanced) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It delivered the best overlap result in the set and one of the strongest medical-jargon transcripts, while preserving the complete developer payload on every input. The speaker count on the crosstalk clip matched the four-person conversation structure. **Where it struggled:** - Its bilingual performance fell off more sharply than AssemblyAI's, with only 12.5% Spanish token recall and a noticeably larger deletion count. It is excellent for hard English audio, but less convincing when the language flips mid-sentence. **What came out:** ![Speechmatics (Ursa/Enhanced) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/b0f36d9c22e145e0b521882a7d8c127e.png?v=1) *Output — The crosstalk request used Speechmatics enhanced mode with diarization enabled and an English language setting.* ![Speechmatics (Ursa/Enhanced) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/6475a5e628f6469bb90133294fa31ca6.png?v=1) *Output — The crosstalk run scored WER 26.63%, with 507 substitutions, 1430 deletions, 81 insertions, and 46.71s latency.* ![Speechmatics (Ursa/Enhanced) output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/d026cc838f184cafa97048efceea118d.png?v=1) *Output — The crosstalk error window showed a long missing span, and the engine over-segmented speakers by returning 5 labels for 4 participants.* ![Speechmatics (Ursa/Enhanced) output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/d987485f2bf34280b1c98566d97f44ab.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 7812 timed tokens in the payload.* ![Speechmatics (Ursa/Enhanced) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/7b2e0ac4924444d39c8bd690f2135982.png?v=1) *Output — The crosstalk job completed as a multi-stage async flow and returned 6230 words from a 7579-word reference.* ![Speechmatics (Ursa/Enhanced) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/881ffe23f38d4c67838ee72daf7f9e2c.png?v=1) *Output — The medical run used the same enhanced configuration, with diarization requested on the single-speaker narration.* ![Speechmatics (Ursa/Enhanced) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/1fcfaa34ccc342b1bfba913ed7d640c5.png?v=1) *Output — The medical run scored WER 3.01% and perfect jargon recall, making it one of the strongest medical results in the set.* ![Speechmatics (Ursa/Enhanced) output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/c0c8e62994434513bd45f97516766135.png?v=1) *Output — The main medical divergence was a small wording error around 'atrophy' versus 'are to afi', not a failure on the domain terms themselves.* ![Speechmatics (Ursa/Enhanced) output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/35765f81b42e4449bfa155a3ddae65d3.png?v=1) *Output — The response showed successful word-level transcription with confidence and a single detected speaker label.* ![Speechmatics (Ursa/Enhanced) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/366ff854b6264f0e9c29e6751d9b9f30.png?v=1) *Output — The medical job completed end to end and returned 2727 words against a 2728-word reference.* ![Speechmatics (Ursa/Enhanced) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/3e05f4976b8a4626b5c94d85bbcd5c9d.png?v=1) *Output — The bilingual run submitted mix_language.mp3 with enhanced mode and speaker diarization enabled.* ![Speechmatics (Ursa/Enhanced) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/54fefc4470df44c393815aa2eb2f0970.png?v=1) *Output — The bilingual run scored WER 25.06%, with 560 substitutions, 981 deletions, 92 insertions, and 12.5% Spanish token recall.* ![Speechmatics (Ursa/Enhanced) output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/76864ef504d34f068d7a4373959e823c.png?v=1) *Output — The bilingual error window dropped the Spanish phrase around 'mi entoces ahora', showing weak multilingual routing on this clip.* ![Speechmatics (Ursa/Enhanced) output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/074b7aeb60c84b3f9b6ff39ef92961a3.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 6907 timed tokens in the payload.* ![Speechmatics (Ursa/Enhanced) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/985619ffd59441038ff310bea3517c28.png?v=1) *Output — The bilingual job completed as a multi-stage async flow and returned 5628 words from a 6517-word reference.* ### ElevenLabs Scribe Very fast and very strong on the medical clip, with good overall accuracy on the hard set, but speaker over-segmentation and middling bilingual performance hold it back. ![ElevenLabs Scribe screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![ElevenLabs Scribe screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![ElevenLabs Scribe screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It was the fastest engine on the bilingual clip and the best engine on the medical clip, while keeping a full developer payload on every input. On clean technical narration it is especially strong. **Where it struggled:** - It over-segmented the crosstalk clip into 5 speakers, and its bilingual transcript was less reliable than the top two engines. The quality is good, but the diarization behavior is not meeting-note clean. **What came out:** ![ElevenLabs Scribe output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/a65fa0dcf9cc40c7a7f0c277499851cd.png?v=1) *Output — The crosstalk request used multipart speech-to-text with diarization and word timestamps enabled.* ![ElevenLabs Scribe output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/f90d0198ceb44a94b003d2fac085336c.png?v=1) *Output — The crosstalk run scored WER 26.67%, with 856 substitutions, 782 deletions, 383 insertions, and 41.76s latency.* ![ElevenLabs Scribe output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/50da30bf2e114e109981ac0053fcb24c.png?v=1) *Output — The crosstalk error window showed a long missing span, and the engine over-segmented speakers by returning 5 labels for 4 participants.* ![ElevenLabs Scribe output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/7d4b12ac00234c99a783c9b0b6829e4c.png?v=1) *Output — The response included word-level timing, logprob confidence, and speaker labels, with 14506 timed tokens in the payload.* ![ElevenLabs Scribe output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/b148be628ad34aee82ff399d58254fd5.png?v=1) *Output — The crosstalk job completed inline and returned 7180 words from a 7579-word reference.* ![ElevenLabs Scribe output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/df92fe9dca8f4f9b880d906b2081150a.png?v=1) *Output — The medical run submitted medical_terms.mp3 with diarization and word timestamps enabled.* ![ElevenLabs Scribe output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/f4ef260e4e2149aebb582b12f606211b.png?v=1) *Output — The medical run scored WER 3.01%, the best medical result in the set, with perfect jargon recall.* ![ElevenLabs Scribe output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/7ac84863d973418484624514c63c6d02.png?v=1) *Output — The medical divergence was a small numeric/wording shift around the reference span, not a missed domain term.* ![ElevenLabs Scribe output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/c0c7fb7bac2d4a2c96a38e5761f93107.png?v=1) *Output — The response showed word-level timing, logprob confidence, and speaker labels, with 5448 timed tokens in the payload.* ![ElevenLabs Scribe output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/ec512ec412d543679f5647f14775de4f.png?v=1) *Output — The medical job completed inline and returned 2743 words from a 2728-word reference.* ![ElevenLabs Scribe output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/47f36a2814894bb8ad2a2ead99b023da.png?v=1) *Output — The bilingual run submitted mix_language.mp3 with diarization and word timestamps enabled.* ![ElevenLabs Scribe output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/03ab3bf8d6284208be137694c1349220.png?v=1) *Output — The bilingual run scored WER 28.57%, with 752 substitutions, 778 deletions, 332 insertions, and 57.5% Spanish token recall.* ![ElevenLabs Scribe output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/6bcf29343e534b19a93c6ea77c584412.png?v=1) *Output — The bilingual error window dropped the Spanish token around 'ahora', but the transcript still retained more of the conversation than many cheaper engines.* ![ElevenLabs Scribe output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/bcf45c2b5c5d45879d72c5c4761a9547.png?v=1) *Output — The response included word-level timing, logprob confidence, and speaker labels, with 12162 timed tokens in the payload.* ![ElevenLabs Scribe output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/808ac747f28a44e0ab00de3e2aa0b3e9.png?v=1) *Output — The bilingual job completed inline and returned 6071 words from a 6517-word reference.* ### AWS Transcribe A steady enterprise async option: good payload, strong diarization, and respectable accuracy, but slower and not as sharp as the top three on the hardest cases. ![AWS Transcribe screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![AWS Transcribe screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![AWS Transcribe screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It returned complete developer payloads on all three inputs and produced one of the stronger bilingual scores outside the top three. It is a stable, broadly usable async engine if you can tolerate slower turnaround. **Where it struggled:** - The overlap clip still lost a lot of content, and the bilingual run retained only 30.0% of the scored Spanish types. It is dependable, but not the most accurate choice here. **What came out:** ![AWS Transcribe output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/91f2c7a206374eb38ca80804ee95db24.png?v=1) *Output — The crosstalk request used the AWS SDK path with speaker labels enabled and S3-backed batch transcription.* ![AWS Transcribe output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/a2b869061df345d3bad937671936f2f3.png?v=1) *Output — The crosstalk run scored WER 33.88%, with 363 substitutions, 2162 deletions, 43 insertions, and 196.83s latency.* ![AWS Transcribe output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/a1770c2bacef406ab10241e356ede8fe.png?v=1) *Output — The crosstalk error window showed a long missing span, but the engine still returned 4 speaker labels for 4 participants.* ![AWS Transcribe output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/034c47c756e8469da08b31740f4c05e4.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 5912 timed tokens in the payload.* ![AWS Transcribe output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/4627f600af174781b3e1fac425f25be0.png?v=1) *Output — The crosstalk job completed through the AWS async flow and returned 5460 words from a 7579-word reference.* ![AWS Transcribe output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/03296f02a67f464f988f85897c04e912.png?v=1) *Output — The medical run submitted medical_terms.mp3 through the same AWS SDK-backed batch path.* ![AWS Transcribe output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/e3414190a65141a98f98591c8fb7428e.png?v=1) *Output — The medical run scored WER 3.63% and perfect jargon recall, but was slower than the top engines.* ![AWS Transcribe output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/365b48bd6ccd46c2acf03e80c2c37dcf.png?v=1) *Output — The medical divergence was a single domain-word miss, where 'cancellous' was rendered as 'cancerous'.* ![AWS Transcribe output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/4a6025224b044895b1da4ef321700e5a.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 2829 timed tokens in the payload.* ![AWS Transcribe output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/6f601bfc316b4337ae5357ff37182c11.png?v=1) *Output — The medical job completed through the AWS async flow and returned 2728 words from a 2728-word reference.* ![AWS Transcribe output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/8fe72d8f7cf64141b6cb43ca72b934ea.png?v=1) *Output — The bilingual run submitted mix_language.mp3 through the same AWS SDK-backed batch path.* ![AWS Transcribe output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/4246c479d33c4f4d9b07a69de5b762c7.png?v=1) *Output — The bilingual run scored WER 23.06%, with 673 substitutions, 665 deletions, 165 insertions, and 30.0% Spanish token recall.* ![AWS Transcribe output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/b201a6a4cda945f0bb3b7f0179e055c2.png?v=1) *Output — The bilingual error window dropped the Spanish phrase around 'mi entoces ahora', but the transcript still kept the conversation more intact than several rivals.* ![AWS Transcribe output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/6674123c061c438fbb1c7e2247b29bbc.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 6311 timed tokens in the payload.* ![AWS Transcribe output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/d35e887bc96c4ffbb937e1072aeafc63.png?v=1) *Output — The bilingual job completed through the AWS async flow and returned 6017 words from a 6517-word reference.* ### Deepgram Nova-3 Fast and reasonably strong on the medical clip, but the hard-case error profile is too uneven for overlap and bilingual audio. ![Deepgram Nova-3 screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![Deepgram Nova-3 screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![Deepgram Nova-3 screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It was very fast, and the medical clip showed that it can handle jargon well when the acoustics are clean. The payload is developer-friendly and structurally complete. **Where it struggled:** - The overlap clip produced too many deletions and insertions, and the bilingual clip was especially weak on Spanish recall. The engine is usable, but its hard-case consistency is not yet good enough to lead the set. **What came out:** ![Deepgram Nova-3 output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/2121f37f2e584a55a39400b76fc91be7.png?v=1) *Output — The crosstalk request used Deepgram Nova-3 with smart formatting, diarization, punctuation, and utterances enabled.* ![Deepgram Nova-3 output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/18d914a4a4ac4fbc90428615fa93c938.png?v=1) *Output — The crosstalk run scored WER 36.27%, with 1096 substitutions, 1143 deletions, 510 insertions, and 63.87s latency.* ![Deepgram Nova-3 output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/3d4527d0398f4d8aae42f0f127fa9996.png?v=1) *Output — The crosstalk divergence replaced a long missing span with an incorrect short phrase, showing a noisy overlap transcript.* ![Deepgram Nova-3 output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/548f2016acaf40a3ae5afa069abbff63.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 13555 timed tokens in the payload.* ![Deepgram Nova-3 output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/37b6598755d040a6ace24f32ae3181fa.png?v=1) *Output — The crosstalk job completed through Deepgram's single-call flow and returned 6946 words from a 7579-word reference.* ![Deepgram Nova-3 output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/d43f90e012bc426ba3ee51c64eab38de.png?v=1) *Output — The medical run used the same Deepgram Nova-3 request shape with diarization and formatting enabled.* ![Deepgram Nova-3 output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/5f1d578b9fb54412a0760ce828998f1d.png?v=1) *Output — The medical run scored WER 5.43%, with perfect jargon recall and the fastest turnaround among the scored inputs.* ![Deepgram Nova-3 output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/ee6a6fb9d3824f7fbccbe2f74c83fa35.png?v=1) *Output — The medical divergence was a missing span in the middle of the anatomy text, but the jargon terms themselves were preserved.* ![Deepgram Nova-3 output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/bd3d89e2e8f04bd7987df34943bb2375.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 5845 timed tokens in the payload.* ![Deepgram Nova-3 output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/edf8364ebee143bfb7590d4ff71bb202.png?v=1) *Output — The medical job completed through Deepgram's single-call flow and returned 2726 words from a 2728-word reference.* ![Deepgram Nova-3 output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/fdc86bd48d4a4edabf5dc1387bd90836.png?v=1) *Output — The bilingual run used the same Deepgram Nova-3 request shape for mix_language.mp3.* ![Deepgram Nova-3 output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/ed4aca0c4fda4186ad44424a20ee6ad0.png?v=1) *Output — The bilingual run scored WER 38.13%, with 793 substitutions, 1259 deletions, 433 insertions, and 3.8% Spanish token recall.* ![Deepgram Nova-3 output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/badb0300a57e4693854946dc20909cd6.png?v=1) *Output — The bilingual divergence dropped the Spanish token 'ahora', and the missed-span list shows very poor code-switch handling.* ![Deepgram Nova-3 output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/a137355f0e084175bd40f8494af395f6.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 11684 timed tokens in the payload.* ![Deepgram Nova-3 output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/9fe6a97ec70043048480624b1def1d6c.png?v=1) *Output — The bilingual job completed through Deepgram's single-call flow and returned 5691 words from a 6517-word reference.* ### Gladia Excellent on jargon and very fast, but the bilingual response duplicated channels and made the longest multilingual transcript unusable without a channel fix. ![Gladia screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![Gladia screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![Gladia screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It was extremely strong on the medical-jargon clip and kept a complete developer payload on every input. The speed is good and the structured output is useful. **Where it struggled:** - The bilingual run duplicated both channels and made the transcript unusable without mono downmixing or explicit channel handling. That is a serious integration issue for mixed-language audio. **What came out:** ![Gladia output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/b0f36d9c22e145e0b521882a7d8c127e.png?v=1) *Output — The crosstalk request used Gladia Solaria with diarization enabled on crosstalk.wav.* ![Gladia output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/6475a5e628f6469bb90133294fa31ca6.png?v=1) *Output — The crosstalk run scored WER 37.35%, with 529 substitutions, 2213 deletions, 89 insertions, and 26.33s latency.* ![Gladia output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/d026cc838f184cafa97048efceea118d.png?v=1) *Output — The crosstalk divergence dropped a long overlap-heavy span, but the engine still returned 4 speaker labels for 4 participants.* ![Gladia output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/d987485f2bf34280b1c98566d97f44ab.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 12968 timed tokens in the payload.* ![Gladia output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/7b2e0ac4924444d39c8bd690f2135982.png?v=1) *Output — The crosstalk job completed through Gladia's multi-stage flow and returned 5455 words from a 7579-word reference.* ![Gladia output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/881ffe23f38d4c67838ee72daf7f9e2c.png?v=1) *Output — The medical run used the same Gladia Solaria request shape with diarization enabled on medical_terms.mp3.* ![Gladia output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/1fcfaa34ccc342b1bfba913ed7d640c5.png?v=1) *Output — The medical run scored WER 4.07% and perfect jargon recall, making it one of the strongest jargon results in the set.* ![Gladia output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/c0c8e62994434513bd45f97516766135.png?v=1) *Output — The medical divergence was a small numeric phrasing difference, not a failure on the domain vocabulary.* ![Gladia output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/35765f81b42e4449bfa155a3ddae65d3.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 5854 timed tokens in the payload.* ![Gladia output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/366ff854b6264f0e9c29e6751d9b9f30.png?v=1) *Output — The medical job completed through Gladia's multi-stage flow and returned 2738 words from a 2728-word reference.* ![Gladia output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/3e05f4976b8a4626b5c94d85bbcd5c9d.png?v=1) *Output — The bilingual run used the same Gladia Solaria request shape on mix_language.mp3.* ![Gladia output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/54fefc4470df44c393815aa2eb2f0970.png?v=1) *Output — The bilingual run scored WER 88.45%, with 1110 substitutions, 203 deletions, 4451 insertions, and 56.2% Spanish token recall.* ![Gladia output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/76864ef504d34f068d7a4373959e823c.png?v=1) *Output — The bilingual divergence showed channel duplication: the transcript repeated phrases and inflated the word count to 10765 words against a 6517-word reference.* ![Gladia output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/074b7aeb60c84b3f9b6ff39ef92961a3.png?v=1) *Output — The response showed two distinct channel values and 26166 timed tokens, which explains the duplicated bilingual transcript.* ![Gladia output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/985619ffd59441038ff310bea3517c28.png?v=1) *Output — The bilingual job completed through Gladia's multi-stage flow and returned 10765 words from a 6517-word reference.* ### Rev AI A conservative async engine with decent overlap handling, moderate multilingual performance, and full developer payloads, but weaker medical-jargon accuracy than the leaders. ![Rev AI screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![Rev AI screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![Rev AI screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It returned full developer payloads and kept its crosstalk transcript closer to the reference than some lower-ranked engines. The output is steady and predictable rather than flashy. **Where it struggled:** - Its medical-jargon run was weaker than the top contenders, and the bilingual clip still lost most of the Spanish tokens. It is usable, but not the most accurate or multilingual-ready engine in this set. **What came out:** ![Rev AI output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/1356772120714e5198dec4b02687c98c.png?v=1) *Output — The crosstalk request used Rev AI's async endpoint with skip_diarization=false.* ![Rev AI output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/464c6b88f7344a558d7400d1d82a2639.png?v=1) *Output — The crosstalk run scored WER 28.33%, with 601 substitutions, 1423 deletions, 123 insertions, and 102.7s latency.* ![Rev AI output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/bd11b2dd8a00413c80b16820854e8a3d.png?v=1) *Output — The crosstalk error window showed a long missing span, and the engine over-segmented speakers into 6 labels for 4 participants.* ![Rev AI output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/062e1ce07ace4bff9622f8770f3977b1.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 6249 timed tokens in the payload.* ![Rev AI output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/6c286541e77f49b7a7982e5bb5691653.png?v=1) *Output — The crosstalk job completed through Rev AI's async flow and returned 6279 words from a 7579-word reference.* ![Rev AI output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/4dabe22d76824fb3927df43bd2ba2993.png?v=1) *Output — The medical run used the same Rev AI async endpoint with diarization disabled at the request level and then reconstructed in output labels.* ![Rev AI output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/bfe3061b7f8843f49a44eca90a32a173.png?v=1) *Output — The medical run scored WER 9.79%, with 195 substitutions, 7 deletions, 65 insertions, and 77.8% jargon recall.* ![Rev AI output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/8b82286c35a149c582eb7bba299a52c5.png?v=1) *Output — The medical divergence missed the term 'cancellous', which was rendered as 'cancerous'.* ![Rev AI output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/4408e2e90f934a60a6320a6ab66f9b53.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 2779 timed tokens in the payload.* ![Rev AI output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/bb3a1ad9948f4125bb1ca2569432010b.png?v=1) *Output — The medical job completed through the Rev AI async flow and returned 2786 words from a 2728-word reference.* ![Rev AI output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/e0a8c47be38547af928bd918e0c885cc.png?v=1) *Output — The bilingual run used the same Rev AI async endpoint for mix_language.mp3.* ![Rev AI output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/c49ce0644577437d9ec2f58817a21fe5.png?v=1) *Output — The bilingual run scored WER 25.16%, with 691 substitutions, 793 deletions, 156 insertions, and 8.8% Spanish token recall.* ![Rev AI output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/b70abab05b054c2584574d32bead1320.png?v=1) *Output — The bilingual error window dropped the Spanish phrase around 'mi entoces ahora', showing weak code-switch handling.* ![Rev AI output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/8e22eb23942646deabaf26df3416ad62.png?v=1) *Output — The response included word-level timing, confidence, and speaker labels, with 5855 timed tokens in the payload.* ![Rev AI output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/dbb7656100ce41d69e44857ee0384cd6.png?v=1) *Output — The bilingual job completed through Rev AI's async flow and returned 5880 words from a 6517-word reference.* ### GroqCloud (Whisper Large-v3) Blazing fast and cheap on accepted files, but the long overlap clip hit the upload cap, so it is risky for long-form transcription work. ![GroqCloud (Whisper Large-v3) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![GroqCloud (Whisper Large-v3) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![GroqCloud (Whisper Large-v3) screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It was by far one of the fastest and cheapest engines on the inputs it accepted, and the medical-jargon score was excellent. On shorter files, the engine is very attractive for throughput and cost. **Where it struggled:** - The longest file hit a hard upload cap, which means this engine is not reliable for the long-form benchmark as configured here. It also returns only segment-level timing and no speaker labels, so it is weaker for downstream captioning and diarization workflows. **What came out:** ![GroqCloud (Whisper Large-v3) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/ab013b63df194e448c11712ab8186bdf.png?v=1) *Output — The crosstalk request used GroqCloud Whisper Large-v3 with a multipart upload and word timestamps enabled.* ![GroqCloud (Whisper Large-v3) output showing Limit evidence](https://cdn.futuresmart.ai/public/aidemos/c7a313c2cd484ee28df2783153b59883.png?v=1) *Output — The API rejected crosstalk.wav because it exceeded the 25 MB upload cap, so no transcript was produced for the longest file.* ![GroqCloud (Whisper Large-v3) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/cc0e6e98581c401db600ef98acc4266a.png?v=1) *Output — The crosstalk run was rejected by the API after 2.59s, with no transcript returned and zero charge.* ![GroqCloud (Whisper Large-v3) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/defb9ac4f0f843e5807d4d5b95cf406a.png?v=1) *Output — The crosstalk job ended in rejected_by_api status, confirming the long-file ceiling.* ![GroqCloud (Whisper Large-v3) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/1c09ace0bf2047c58af8bba1789cfee4.png?v=1) *Output — The medical run used the same GroqCloud Whisper Large-v3 request shape on medical_terms.mp3.* ![GroqCloud (Whisper Large-v3) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/daf30c9932cb40fe8e6272d7e228ded0.png?v=1) *Output — The medical run scored WER 3.15%, with perfect jargon recall and the fastest turnaround in the set.* ![GroqCloud (Whisper Large-v3) output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/8e7131164ce34c45b3a898cd129b8d78.png?v=1) *Output — The medical divergence was a short omitted phrase, not a vocabulary failure.* ![GroqCloud (Whisper Large-v3) output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/ffec0f6b76244d14973ab50e12df5b0e.png?v=1) *Output — The response showed a successful transcription with segment-level timing and no speaker labels.* ![GroqCloud (Whisper Large-v3) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/c163a9e1af97437f9a22ed13c46a2e6e.png?v=1) *Output — The medical job completed through the single-call flow and returned 2730 words from a 2728-word reference.* ![GroqCloud (Whisper Large-v3) output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/38b4bb13583f40b690a8c4962bfbb4ce.png?v=1) *Output — The bilingual run used the same GroqCloud Whisper Large-v3 request shape on mix_language.mp3.* ![GroqCloud (Whisper Large-v3) output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/6d6c951fe0314750b151a61c62461e40.png?v=1) *Output — The bilingual run scored WER 27.94%, with 636 substitutions, 1052 deletions, 133 insertions, and 53.8% Spanish token recall.* ![GroqCloud (Whisper Large-v3) output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/d22a006544b443cbaff758ce3756ce8b.png?v=1) *Output — The bilingual divergence dropped a Spanish phrase and showed a noisy English-only rewrite.* ![GroqCloud (Whisper Large-v3) output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/db416f394c174961beb7872d7e51419e.png?v=1) *Output — The response showed a successful transcription with segment-level timing and no speaker labels.* ![GroqCloud (Whisper Large-v3) output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/594567a6b59247af9e6f7b787fbbe51e.png?v=1) *Output — The bilingual job completed through the single-call flow and returned 5598 words from a 6517-word reference.* ### OpenAI Simple transcription API, but the long overlap file hit the 25 MB limit and the developer payload is thinner than the best options. ![OpenAI screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![OpenAI screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![OpenAI screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It returned usable transcripts for the shorter files and did reasonably well on the medical clip. The API is straightforward to integrate when the input stays under the limit. **Where it struggled:** - The longest meeting clip was rejected outright, and the payload lacks confidence and speaker labels in the scored runs we have here. For this benchmark, that makes it a weaker fit than the top engines. **What came out:** ![OpenAI output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/b0f36d9c22e145e0b521882a7d8c127e.png?v=1) *Output — The crosstalk request used the OpenAI transcription endpoint with gpt-4o-transcribe and word timestamps enabled.* ![OpenAI output showing Limit evidence](https://cdn.futuresmart.ai/public/aidemos/c7a313c2cd484ee28df2783153b59883.png?v=1) *Output — The API rejected crosstalk.wav because it exceeded the 25 MB upload cap, so no transcript was produced for the longest file.* ![OpenAI output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/6475a5e628f6469bb90133294fa31ca6.png?v=1) *Output — The crosstalk run was rejected by the API after 50.08s, with no transcript returned and zero charge.* ![OpenAI output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/7b2e0ac4924444d39c8bd690f2135982.png?v=1) *Output — The crosstalk job ended in rejected_by_api status, confirming the hard file-size ceiling.* ![OpenAI output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/881ffe23f38d4c67838ee72daf7f9e2c.png?v=1) *Output — The medical run used the same OpenAI transcription endpoint and stayed under the file-size cap.* ![OpenAI output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/35765f81b42e4449bfa155a3ddae65d3.png?v=1) *Output — The response began with successful transcription text and word-level timestamps, but no confidence or speaker labels were present.* ![OpenAI output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/1fcfaa34ccc342b1bfba913ed7d640c5.png?v=1) *Output — The medical run scored WER 5.17%, with 54 substitutions, 76 deletions, 11 insertions, and 88.9% jargon recall.* ![OpenAI output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/c0c8e62994434513bd45f97516766135.png?v=1) *Output — The medical divergence missed the domain term 'trabeculae', which was rendered as 'trabeculi'.* ![OpenAI output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/366ff854b6264f0e9c29e6751d9b9f30.png?v=1) *Output — The medical job completed inline and returned 2663 words from a 2728-word reference.* ![OpenAI output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/3e05f4976b8a4626b5c94d85bbcd5c9d.png?v=1) *Output — The bilingual run used the same OpenAI transcription endpoint on mix_language.mp3.* ![OpenAI output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/074b7aeb60c84b3f9b6ff39ef92961a3.png?v=1) *Output — The response began with successful English transcription text and word-level timestamps, but no confidence or speaker labels were present.* ![OpenAI output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/54fefc4470df44c393815aa2eb2f0970.png?v=1) *Output — The bilingual run scored WER 26.12%, with 623 substitutions, 947 deletions, 132 insertions, and 46.2% Spanish recall.* ![OpenAI output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/76864ef504d34f068d7a4373959e823c.png?v=1) *Output — The bilingual divergence dropped the Spanish phrase around 'mi entonces ahora', showing weak code-switch handling.* ![OpenAI output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/985619ffd59441038ff310bea3517c28.png?v=1) *Output — The bilingual job completed inline and returned 5702 words from a 6517-word reference.* ### Google Cloud STT v2 Returned transcripts, but latency was very high and the hard-audio accuracy was well behind the best options, with no useful diarization here. ![Google Cloud STT v2 screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1) *Screenshot — 35:43 overlapping meeting audio from the AMI corpus, used to test crosstalk robustness and diarization.* ![Google Cloud STT v2 screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1) *Screenshot — 18:44 medical narration from Gray's Anatomy, used to test technical vocabulary and proper-noun accuracy.* ![Google Cloud STT v2 screenshot showing Audio input](https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1) *Screenshot — 32:18 Spanish-English code-switching conversation from the Bangor Miami corpus, used to test multilingual routing.* **What worked:** - It could still preserve some medical jargon, and the API did return transcripts rather than failing completely. For a general cloud STT service, that is not nothing. **Where it struggled:** - Latency was very high, hard-audio accuracy was the weakest in the set, and the diarization path did not work in this configuration. This is not the engine to choose for hard meeting audio or mixed-language speech. **What came out:** ![Google Cloud STT v2 output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/5d1a045b44a84036a74c5c845b34857d.png?v=1) *Output — The crosstalk request used Google Cloud STT v2 with chirp settings, word offsets, and diarization requested.* ![Google Cloud STT v2 output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/e9a7258cbfd349f99481b2f388535c5e.png?v=1) *Output — The response began with a transcript snippet, but the payload did not return usable speaker labels for diarization.* ![Google Cloud STT v2 output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/77d878df5d1e426985af5eac4e7e7078.png?v=1) *Output — The crosstalk run scored WER 43.50%, with 633 substitutions, 2614 deletions, 50 insertions, and 810.6s latency.* ![Google Cloud STT v2 output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/227c888479f74338800ddd26a6edbf4b.png?v=1) *Output — The crosstalk divergence replaced a long overlap-heavy span with a short incorrect rewrite, and diarization was absent.* ![Google Cloud STT v2 output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/1c22860f93a24224ac73808f217c9aae.png?v=1) *Output — The crosstalk batch job completed, but the trace notes that diarization was rejected by the API in the final request.* ![Google Cloud STT v2 output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/ebaad9a06a5241278425d7c57c92fea8.png?v=1) *Output — The medical run used the same Google Cloud STT v2 request shape with word offsets and diarization requested.* ![Google Cloud STT v2 output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/3aea11164a424836b02dc200b4afd353.png?v=1) *Output — The response began with a short transcript snippet, but the payload still did not return usable speaker labels.* ![Google Cloud STT v2 output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/1b0bb00e62964c42852b8bcb344195cd.png?v=1) *Output — The medical run scored WER 13.09%, with 228 substitutions, 49 deletions, 80 insertions, and 100% jargon recall.* ![Google Cloud STT v2 output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/571bbbb305cb44be8d122e742752aa26.png?v=1) *Output — The medical divergence omitted a long middle span, showing that the engine could preserve jargon but still miss large chunks of text.* ![Google Cloud STT v2 output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/8d4541f5b9ca449ab0127450a7f7f656.png?v=1) *Output — The medical batch job completed, but the trace again notes that diarization was rejected by the API.* ![Google Cloud STT v2 output showing Request configuration](https://cdn.futuresmart.ai/public/aidemos/7060bb160de946c089a2f57109007c82.png?v=1) *Output — The bilingual run used the same Google Cloud STT v2 request shape with word offsets and diarization requested.* ![Google Cloud STT v2 output showing Raw API response](https://cdn.futuresmart.ai/public/aidemos/37ce421c5edd4275b80441fc6ce9f7e6.png?v=1) *Output — The response began with a short bilingual transcript snippet, but the payload did not return usable speaker labels.* ![Google Cloud STT v2 output showing Run metrics](https://cdn.futuresmart.ai/public/aidemos/b62c1aa554c54c0e8435094009f9dc1a.png?v=1) *Output — The bilingual run scored WER 56.13%, with 716 substitutions, 2884 deletions, 58 insertions, and 5.0% Spanish recall.* ![Google Cloud STT v2 output showing Transcript detail](https://cdn.futuresmart.ai/public/aidemos/843dddff2af74fdfb4bc61e4655bb5c5.png?v=1) *Output — The bilingual divergence dropped the Spanish phrase around 'mi entonces ahora', showing very poor code-switch handling.* ![Google Cloud STT v2 output showing Execution trace](https://cdn.futuresmart.ai/public/aidemos/c61bace5debf49568670253087978efb.png?v=1) *Output — The bilingual batch job completed, but the trace notes that the final request still had unsupported diarization fields.* ## Evidence (first-party, tested) *72 tested cells · 61/72 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:aws-transcribe·cross·automation-level`.* | Tool | Criterion | Scenario | Verdict | Score | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | --- | | AWS Transcribe | Automation level | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3d810cf7afca4101b2648301b88e0dce.png?v=1) | `ev:aws-transcribe·cross·automation-level` | | AWS Transcribe | Automation level | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/47be998587d84e5ba0c33ef977d1569d.png?v=1) | `ev:aws-transcribe·overlapping-meeting-speech-with-cross-talk·automation-level` | | AWS Transcribe | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a8952ff0bc8a4c7ebadaf9ec2df44ca9.png?v=1) | `ev:aws-transcribe·bilingual-spanish-english-code-switching-speech·automation-level` | | AWS Transcribe | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/652e08c8c4b64c588cd4f3a90c91bf9f.png?v=1) | `ev:aws-transcribe·medical-anatomy-narration-with-dense-jargon·automation-level` | | AWS Transcribe | Export | cross-scenario | ✓ worked | — | 👁 observed | `ev:aws-transcribe·cross·export` | | AWS Transcribe | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7ac7fc79af1d45e98ed0b1a23f62a9ba.png?v=1) | `ev:aws-transcribe·medical-anatomy-narration-with-dense-jargon·export` | | AWS Transcribe | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f7e03bf80a8e4884aef9515ac9d05739.png?v=1) | `ev:aws-transcribe·overlapping-meeting-speech-with-cross-talk·export` | | AWS Transcribe | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3d810cf7afca4101b2648301b88e0dce.png?v=1) | `ev:aws-transcribe·overlapping-meeting-speech-with-cross-talk·input-handling` | | AWS Transcribe | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-aws-transcribe-a08a2f3297a5.md) | `ev:aws-transcribe·cross·input-handling` | | AWS Transcribe | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/439691e323bc456a84344fab78c75c30.png?v=1) | `ev:aws-transcribe·bilingual-spanish-english-code-switching-speech·input-handling` | | AWS Transcribe | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/892fcc8e770c4aa5a0802bc2cb1eaeee.png?v=1) | `ev:aws-transcribe·medical-anatomy-narration-with-dense-jargon·input-handling` | | AWS Transcribe | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 3.6/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/892fcc8e770c4aa5a0802bc2cb1eaeee.png?v=1) | `ev:aws-transcribe·medical-anatomy-narration-with-dense-jargon·output-quality` | | AWS Transcribe | Output quality | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f7e03bf80a8e4884aef9515ac9d05739.png?v=1) | `ev:aws-transcribe·cross·output-quality` | | AWS Transcribe | Output quality | Overlapping meeting speech with cross-talk | ⚠ struggled | 33.9/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3d810cf7afca4101b2648301b88e0dce.png?v=1) | `ev:aws-transcribe·overlapping-meeting-speech-with-cross-talk·output-quality` | | AWS Transcribe | Output quality | Bilingual Spanish-English code-switching speech | ◐ mixed | 23.1/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/439691e323bc456a84344fab78c75c30.png?v=1) | `ev:aws-transcribe·bilingual-spanish-english-code-switching-speech·output-quality` | | ElevenLabs Scribe | Automation level | cross-scenario | ◐ mixed | — | 👁 observed | `ev:elevenlabs-scribe·cross·automation-level` | | ElevenLabs Scribe | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/0ad15b39196645478e4b8c12251fd5f5.png?v=1) | `ev:elevenlabs-scribe·medical-anatomy-narration-with-dense-jargon·automation-level` | | ElevenLabs Scribe | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/2099044eff544dfa915b029d4dab1245.png?v=1) | `ev:elevenlabs-scribe·bilingual-spanish-english-code-switching-speech·automation-level` | | ElevenLabs Scribe | Automation level | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/b3e5ca72a9b7408daddd0729d4c817d8.png?v=1) | `ev:elevenlabs-scribe·overlapping-meeting-speech-with-cross-talk·automation-level` | | ElevenLabs Scribe | Export | cross-scenario | ✓ worked | — | 👁 observed | `ev:elevenlabs-scribe·cross·export` | | ElevenLabs Scribe | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/666aae8148fd4eea870ff63c3327d3cb.png?v=1) | `ev:elevenlabs-scribe·overlapping-meeting-speech-with-cross-talk·input-handling` | | ElevenLabs Scribe | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/0f370484b7634847b7a0b078a7a3082e.png?v=1) | `ev:elevenlabs-scribe·bilingual-spanish-english-code-switching-speech·input-handling` | | ElevenLabs Scribe | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3fc6bef146c8454bb016e0f12672c04c.png?v=1) | `ev:elevenlabs-scribe·medical-anatomy-narration-with-dense-jargon·input-handling` | | ElevenLabs Scribe | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/e44c577a896a41ed81b79e2986cb72f3.png?v=1) | `ev:elevenlabs-scribe·cross·input-handling` | | ElevenLabs Scribe | Output quality | Overlapping meeting speech with cross-talk | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/e44c577a896a41ed81b79e2986cb72f3.png?v=1) | `ev:elevenlabs-scribe·overlapping-meeting-speech-with-cross-talk·output-quality` | | ElevenLabs Scribe | Output quality | Bilingual Spanish-English code-switching speech | ⚠ struggled | 28.6/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/1cdb5610e7af4f10abf1e8995c3178f0.png?v=1) | `ev:elevenlabs-scribe·bilingual-spanish-english-code-switching-speech·output-quality` | | ElevenLabs Scribe | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 3.0/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/799fe3740fed4fad962f4d443f4e44d2.png?v=1) | `ev:elevenlabs-scribe·medical-anatomy-narration-with-dense-jargon·output-quality` | | ElevenLabs Scribe | Output quality | cross-scenario | ◐ mixed | — | 👁 observed | `ev:elevenlabs-scribe·cross·output-quality` | | Gladia | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f84aa813e76242f58daae74ae5190208.png?v=1) | `ev:gladia·bilingual-spanish-english-code-switching-speech·automation-level` | | Gladia | Automation level | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5e57637448d1453dbfcfce4f0799a344.png?v=1) | `ev:gladia·overlapping-meeting-speech-with-cross-talk·automation-level` | | Gladia | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/fce72a78e0cc44989d44de66037f8d9f.png?v=1) | `ev:gladia·medical-anatomy-narration-with-dense-jargon·automation-level` | | Gladia | Automation level | cross-scenario | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5a2cf23b4acd4b16ba60dce66a8dfd58.mp4?v=1) | `ev:gladia·cross·automation-level` | | Gladia | Export | cross-scenario | ✓ worked | — | 👁 observed | `ev:gladia·cross·export` | | Gladia | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1) | `ev:gladia·bilingual-spanish-english-code-switching-speech·export` | | Gladia | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1) | `ev:gladia·medical-anatomy-narration-with-dense-jargon·export` | | Gladia | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1) | `ev:gladia·overlapping-meeting-speech-with-cross-talk·export` | | Gladia | Input handling | cross-scenario | ✓ worked | — | 👁 observed | `ev:gladia·cross·input-handling` | | Gladia | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5c7b1290c22543ef8e959f985700ea8c.png?v=1) | `ev:gladia·bilingual-spanish-english-code-switching-speech·input-handling` | | Gladia | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/631c69b8537241debdaf55985e3edd25.png?v=1) | `ev:gladia·overlapping-meeting-speech-with-cross-talk·input-handling` | | Gladia | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4d8a73887cd543419cd93d3f9a87ad20.png?v=1) | `ev:gladia·medical-anatomy-narration-with-dense-jargon·input-handling` | | Gladia | Output quality | cross-scenario | ◐ mixed | — | 👁 observed | `ev:gladia·cross·output-quality` | | Gladia | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 4.1/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1) | `ev:gladia·medical-anatomy-narration-with-dense-jargon·output-quality` | | Gladia | Output quality | Bilingual Spanish-English code-switching speech | ✗ failed | 88.5/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1) | `ev:gladia·bilingual-spanish-english-code-switching-speech·output-quality` | | Gladia | Output quality | Overlapping meeting speech with cross-talk | ⚠ struggled | 37.4/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1) | `ev:gladia·overlapping-meeting-speech-with-cross-talk·output-quality` | | OpenAI | Automation level | cross-scenario | ◐ mixed | — | 👁 observed | `ev:openai·cross·automation-level` | | OpenAI | Automation level | Overlapping meeting speech with cross-talk | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5d977dd59cdd462a872dd6cb66d22d78.png?v=1) | `ev:openai·overlapping-meeting-speech-with-cross-talk·automation-level` | | OpenAI | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/82b631a46e6d4fe69367848d41e4ab30.png?v=1) | `ev:openai·bilingual-spanish-english-code-switching-speech·automation-level` | | OpenAI | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a815e76234284ecb939eb0494ec803c2.png?v=1) | `ev:openai·medical-anatomy-narration-with-dense-jargon·automation-level` | | OpenAI | Input handling | cross-scenario | ◐ mixed | — | 👁 observed | `ev:openai·cross·input-handling` | | OpenAI | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/b5785d77b36c452491f53ad042d08933.png?v=1) | `ev:openai·bilingual-spanish-english-code-switching-speech·input-handling` | | OpenAI | Input handling | Overlapping meeting speech with cross-talk | ✗ failed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/fbd80f8132ee419f945a3b2c69f3a11a.png?v=1) | `ev:openai·overlapping-meeting-speech-with-cross-talk·input-handling` | | OpenAI | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7697f11bfd184f5687341d858ade6ca8.png?v=1) | `ev:openai·medical-anatomy-narration-with-dense-jargon·input-handling` | | OpenAI | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 5.2/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7078260079434af596db178f2472cfde.png?v=1) | `ev:openai·medical-anatomy-narration-with-dense-jargon·output-quality` | | OpenAI | Output quality | Bilingual Spanish-English code-switching speech | ⚠ struggled | 26.1/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3648a6bf5f8f440a808e7c18162a6476.png?v=1) | `ev:openai·bilingual-spanish-english-code-switching-speech·output-quality` | | OpenAI | Output quality | cross-scenario | ◐ mixed | — | 👁 observed | `ev:openai·cross·output-quality` | | OpenAI | Output quality | Overlapping meeting speech with cross-talk | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/bf9624af2d0e4a3a85d973f1ffb348a1.png?v=1) | `ev:openai·overlapping-meeting-speech-with-cross-talk·output-quality` | | Rev AI | Automation level | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-rev-ai-57b888ecc885.md) | `ev:rev-ai·cross·automation-level` | | Rev AI | Automation level | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3f3f9d86a4e24b28b6e8b5c05c51e846.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·automation-level` | | Rev AI | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a8441d00948d4bcfb4c569ee592ed81e.png?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·automation-level` | | Rev AI | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5ae09ce9f486423cbca1fb334cd66af7.png?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·automation-level` | | Rev AI | Export | cross-scenario | ✓ worked | — | 👁 observed | `ev:rev-ai·cross·export` | | Rev AI | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·export` | | Rev AI | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·export` | | Rev AI | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·export` | | Rev AI | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-rev-ai-57b888ecc885.md) | `ev:rev-ai·cross·input-handling` | | Rev AI | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/012cfa7be3e94f558e8b420dd26cb8a1.png?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·input-handling` | | Rev AI | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a777b5de6fea4b6693c6e9ba2e5cde69.png?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·input-handling` | | Rev AI | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/50f3d66e58ca4d0abecc6b021fa45e62.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·input-handling` | | Rev AI | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 9.8/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·output-quality` | | Rev AI | Output quality | Overlapping meeting speech with cross-talk | ⚠ struggled | 28.3/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·output-quality` | | Rev AI | Output quality | cross-scenario | ◐ mixed | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-rev-ai-57b888ecc885.md) | `ev:rev-ai·cross·output-quality` | | Rev AI | Output quality | Bilingual Spanish-English code-switching speech | ⚠ struggled | 25.2/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·output-quality` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. ## Final Take AssemblyAI (Universal) is the page’s winner, and that has to stay the headline: it ties the top group on the only deciding check, Output quality at 3.0/5.0, while also scoring 5.0/5.0 on Automation level, Export, and Input handling. The main trade-off is that it does not separate itself on output quality from Speechmatics, ElevenLabs Scribe, AWS Transcribe, GroqCloud, or OpenAI, so the win here is about the strongest all-around measured package rather than a higher transcription score. Speechmatics is the closest alternative in the same output-quality tier, with nearly the same workflow scores. ElevenLabs Scribe stands out for dense medical narration, AWS Transcribe for reliable batch transcription with rich payloads. GroqCloud and OpenAI are also in the 3.0/5.0 output tier, but their weaker input/export limits make them less balanced. Below that, Deepgram Nova-3, Gladia, Rev AI, and Google Cloud STT v2 all fall to 2.0/5.0 on output quality, so they are better viewed as workflow-oriented options with accuracy caveats, not top picks for this ranking. Tested as of 2026-09-02 · re-verified monthly. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text transcription, speaker diarization, or audio transcription system for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).