--- title: "Gladia" type: "AI Tool" url: "https://aidemos.com/tools/gladia" description: "We fed Gladia jargon-heavy and crosstalk audio: it returned word-level timestamps, confidence, and speaker labels, but overlap handling broke spans." category: "audio-speech" website: "https://docs.gladia.io/" published: "2026-08-13T09:18:22.552309+00:00" updated: "2026-09-01T02:58:51.527389+00:00" evidenceCount: 16 verifiedCount: 15 coverage: "dense" --- # Gladia Batch STT with rich word-level metadata and strong jargon recall, but overlap handling is weak and bilingual WER needs a mono-downmixed rerun. ## TL;DR Verdict **Strong metadata and jargon performance, but not uniformly reliable** **Where it wins:** - You need a batch STT API that returns word-level timestamps, confidence, and speaker labels. - You care about technical jargon recall on dense domain audio. - You can rerun stereo or code-switching audio with mono downmix or explicit channel control before trusting the bilingual score. **Main limitation:** You need reliable overlap handling on crosstalk-heavy audio. **Pricing:** Starter (pay-as-you-go) Async $0.61/hr · Real-time $0.75/hr · Growth (annual commitment) Async as low as $0.20/hr · Real-time as low as $0.25/hr · Enterprise Custom — contact sales `Batch API` · `Word-level metadata` · `Jargon recall 100%` · `Channel-duplication caveat` **Website:** [Visit Gladia](https://docs.gladia.io/) ## Evidence (first-party, tested) *16 tested cells · 15/16 artifact-verified. Cite a cell by its Evidence ID, e.g. `ev:gladia·overlapping-meeting-speech-with-cross-talk·automation-level`.* | Criterion | Scenario | Verdict | Proof | Evidence ID | | --- | --- | --- | --- | --- | | Automation level | Overlapping meeting speech with cross-talk | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5e57637448d1453dbfcfce4f0799a344.png?v=1) | `ev:gladia·overlapping-meeting-speech-with-cross-talk·automation-level` | | Automation level | Medical anatomy narration with dense jargon | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/fce72a78e0cc44989d44de66037f8d9f.png?v=1) | `ev:gladia·medical-anatomy-narration-with-dense-jargon·automation-level` | | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f84aa813e76242f58daae74ae5190208.png?v=1) | `ev:gladia·bilingual-spanish-english-code-switching-speech·automation-level` | | Automation level | cross-scenario | ◐ mixed | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5a2cf23b4acd4b16ba60dce66a8dfd58.mp4?v=1) | `ev:gladia·cross·automation-level` | | Export | Medical anatomy narration with dense jargon | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/d16009ea62c2482c90da150b294bffaf.png?v=1) | `ev:gladia·medical-anatomy-narration-with-dense-jargon·export` | | Export | Overlapping meeting speech with cross-talk | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/1e38730bc27c4fa4a44124570255061f.png?v=1) | `ev:gladia·overlapping-meeting-speech-with-cross-talk·export` | | Export | Bilingual Spanish-English code-switching speech | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/83d626d9328142088936ce4690dfc7a3.png?v=1) | `ev:gladia·bilingual-spanish-english-code-switching-speech·export` | | Export | cross-scenario | ✓ worked | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-gladia-bc37c24594b6.md) | `ev:gladia·cross·export` | | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5c7b1290c22543ef8e959f985700ea8c.png?v=1) | `ev:gladia·bilingual-spanish-english-code-switching-speech·input-handling` | | Input handling | cross-scenario | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/abce2c2540df43dda26364746aeaca76.jpeg?v=1) | `ev:gladia·cross·input-handling` | | Input handling | Medical anatomy narration with dense jargon | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4d8a73887cd543419cd93d3f9a87ad20.png?v=1) | `ev:gladia·medical-anatomy-narration-with-dense-jargon·input-handling` | | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/631c69b8537241debdaf55985e3edd25.png?v=1) | `ev:gladia·overlapping-meeting-speech-with-cross-talk·input-handling` | | Output quality | cross-scenario | ◐ mixed | 👁 observed | `ev:gladia·cross·output-quality` | | Output quality | Overlapping meeting speech with cross-talk | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/631c69b8537241debdaf55985e3edd25.png?v=1) | `ev:gladia·overlapping-meeting-speech-with-cross-talk·output-quality` | | Output quality | Bilingual Spanish-English code-switching speech | ◐ mixed | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5c7b1290c22543ef8e959f985700ea8c.png?v=1) | `ev:gladia·bilingual-spanish-english-code-switching-speech·output-quality` | | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4d8a73887cd543419cd93d3f9a87ad20.png?v=1) | `ev:gladia·medical-anatomy-narration-with-dense-jargon·output-quality` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. > **Strong metadata and jargon performance, but not uniformly reliable** > > Gladia looks strong as a batch STT API for developer workflows: every scored run returned word-level timestamps, confidence, and speaker labels, and the medical-jargon sample scored very well with 100% jargon recall. But the crosstalk case missed large spans, and the bilingual run is confounded by channel duplication, so that WER should not be treated as a clean accuracy result until it is rerun with mono downmix or explicit channel control. ## Demo Recording [Video: Gladia demo recording](https://cdn.futuresmart.ai/public/aidemos/f7e3ebe05a4a4f4bb9f01caaac8d611d.mp4?v=1) *Video — Screen-recording workflow showing the Gladia benchmark setup in Finder, Terminal, and RStudio.* ## Feature-by-Feature Breakdown ### Batch Transcription Completes pre-recorded audio as batch transcription jobs end to end. The capability was exercised on crosstalk.wav, medical_terms.mp3, and mix_language.mp3, with job completion, latency, cost, and transcript outputs observed across those runs. **Input:** > **Image** **Output:** **Input:** > **Image** **Output:** **Input:** > **Image** **Output:** **Bottom line:** The batch workflow itself is solid: all three jobs completed with status scored, and the benchmark captured latency, cost, and returned-word counts for comparison. ### Structured Transcript Output **Verdict:** Consistent metadata export across all three inputs. Returns machine-readable transcript payloads with word-level timing, confidence scores, speaker labels, speaker counts, and related metadata. Across the tested audio, the responses were deep JSON structures that were easy to consume downstream. **Input:** > **Image** **Output:** **Input:** > **Image** **Output:** **Input:** > **Image** **Output:** **Bottom line:** The output shape is stable and integration-friendly: word timestamps, confidence, and speaker labels are always present in the scored runs. ### Robust Speech Transcription **Verdict:** Weak on crosstalk-heavy audio. Handles difficult speech conditions such as overlapping speakers, domain-specific jargon, and Spanish-English code-switching. The capability was exercised on crosstalk-heavy meeting audio, a medical lecture, bilingual files, and speaker-labeled overlap variants. **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Bottom line:** It recognized the number of speakers, but overlap-heavy speech still produced a badly incomplete transcript. ## Official pricing Async and real-time rates differ, and the paid plans bundle the core speech features used in this benchmark. | Plan | Price | Notes | | --- | --- | --- | | Starter (pay-as-you-go) | Async $0.61/hr · Real-time $0.75/hr | €50 free credits one-time; 25 async concurrent requests; 30 real-time concurrent; no uptime SLA or priority queue. | | Growth (annual commitment) | Async as low as $0.20/hr · Real-time as low as $0.25/hr | Flexible concurrency; custom volume discounts; 99.9% uptime SLA; priority queue; model-training opt-out. | | Enterprise | Custom — contact sales | Unlimited concurrency; zero data retention; SLAs; custom hosting; custom models and fine-tuning. | *Source: https://www.gladia.io/pricing, accessed 2026-08-14. The report notes that paid plans include diarization, automatic language detection/switching, word-level timestamps, and 100+ languages; Enterprise adds zero data retention and custom hosting.* ## Is It Right For You? **Use it if** - You need a batch STT API that returns word-level timestamps, confidence, and speaker labels. - You care about technical jargon recall on dense domain audio. - You can rerun stereo or code-switching audio with mono downmix or explicit channel control before trusting the bilingual score. **Skip it if** - You need reliable overlap handling on crosstalk-heavy audio. - You need a trustworthy bilingual WER without first rerunning channel-controlled audio. - You need one of the cheapest options; this benchmark ranked Gladia 9/10 on price. ## Classification - **Category:** audio-speech - **Subcategory:** audio-to-text - **Type:** speech - **Built for:** Other ## Frequently Asked Questions **Q: Does Gladia return word timestamps, confidence scores, and speaker labels?** Yes. Every scored run showed a stable developer payload with word_timestamps, confidence, and speaker_labels present, and the raw responses reported payload depth 3/3. **Q: How accurate was Gladia on medical jargon?** Very good. On the Gray's Anatomy sample it scored 4.07% WER, returned 2,738 words against a 2,728-word reference, and achieved 100.0% recall on the scored jargon terms. **Q: How did Gladia do on overlapping speech?** Poorly on transcript completeness. On the crosstalk-heavy AMI clip it scored 37.35% WER with 2,213 deletions, though it did detect 4 speakers, matching the true participant count. **Q: Is the bilingual WER trustworthy?** No, not as recorded here. The bilingual run carried 2 distinct channel values and doubled the transcript length, so the 88.45% WER is an artefact of channel duplication. It should be rerun with mono downmix or explicit channel control before being cited. **Q: What pricing did Gladia list?** The vendor pricing page lists Starter at async $0.61/hr and real-time $0.75/hr, Growth as low as async $0.20/hr and real-time $0.25/hr, and Enterprise as custom pricing. **Q: Did this benchmark measure streaming latency?** No. This was a batch benchmark, and the report says per-call timings for the multi-stage HTTP flow were not instrumented when the run executed. ## Similar Tools AI tools similar to Gladia: - [AssemblyAI](https://aidemos.com/tools/assemblyai-speech-to-text) — Fast batch speech-to-text with rich metadata, strong jargon and mixed-language results, but overlap-heavy meetings can still lose too much. - [Speechmatics](https://aidemos.com/tools/speechmatics) — Strong batch STT for hard English audio, but weak on code-switching as configured. - [Google Cloud Speech-to-Text](https://aidemos.com/tools/google-cloud-speech-to-text) — Timed batch transcripts for mostly English, jargon-heavy audio — but not for diarization or code-switching. - [AWS Transcribe](https://aidemos.com/tools/aws-transcribe) — Batch speech-to-text with word-level metadata, but accuracy drops on overlap and code-switching. - [ElevenLabs Scribe](https://aidemos.com/tools/elevenlabs-scribe) — Fast batch speech-to-text with word-level metadata, strongest on jargon and weaker on overlap/code-switching. - [Rev AI](https://aidemos.com/tools/rev-ai) — Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on crosstalk and code-switching. - [Deepgram](https://aidemos.com/tools/deepgram) — Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching. - [OpenAI](https://aidemos.com/tools/openai) — Batch speech-to-text with word timestamps, but a strict upload cap and weak multilingual performance make it a mixed fit for hard audio. - [ElevenLabs](https://aidemos.com/tools/elevenlabs) — Natural-sounding voice cloning and narration, but with only approximate voice identity. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text transcription, audio transcription, or batch transcription pipeline for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).