--- title: "Deepgram" type: "AI Tool" url: "https://aidemos.com/tools/deepgram" description: "We tested batch speech-to-text with word-level metadata and speaker labels; medical jargon stayed accurate, but crosstalk and code-switching spiked WER." category: "audio-speech" website: "https://deepgram.com" published: "2026-08-20T12:32:56.034069+00:00" updated: "2026-09-04T11:15:26.570114+00:00" evidenceCount: 16 verifiedCount: 14 coverage: "dense" --- # Deepgram Batch speech-to-text with word-level metadata and speaker labels, but weak on crosstalk and code-switching. ## TL;DR Verdict **Good batch STT metadata, uneven hard-case accuracy** **Where it wins:** - you need a batch STT API that returns word-level timing, confidence, and speaker labels - you mainly transcribe single-language technical narration and care about jargon recall - you care more about batch throughput than live streaming **Main limitation:** you need reliable overlapping-speech handling for crosstalk-heavy meetings **Pricing:** Pay As You Go $200 free credit, then pay-as-you-go · Growth $4,000+ / year prepaid credits · Enterprise Requires sales contact `Word timestamps` · `Speaker labels` · `Jargon recall 100%` · `Spanish recall 3.8%` **Website:** [Visit Deepgram](https://deepgram.com) ## Evidence (first-party, tested) *16 tested cells · 14/16 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:deepgram·cross·automation-level`.* | Criterion | Scenario | Verdict | Score | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | | Automation level | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/9e14466f3c7342fb8bf561b1cf8f7043.png?v=1) | `ev:deepgram·cross·automation-level` | | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3f46ad25af044fcea47435b057d24056.png?v=1) | `ev:deepgram·medical-anatomy-narration-with-dense-jargon·automation-level` | | Automation level | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/84e7dcb47c124944ad4547c508cf5ee5.png?v=1) | `ev:deepgram·overlapping-meeting-speech-with-cross-talk·automation-level` | | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/353dc535361142fbab8920650fb7d1ca.png?v=1) | `ev:deepgram·bilingual-spanish-english-code-switching-speech·automation-level` | | Export | cross-scenario | ✓ worked | — | 👁 observed | `ev:deepgram·cross·export` | | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1) | `ev:deepgram·medical-anatomy-narration-with-dense-jargon·export` | | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1) | `ev:deepgram·bilingual-spanish-english-code-switching-speech·export` | | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1) | `ev:deepgram·overlapping-meeting-speech-with-cross-talk·export` | | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/eb06e28bd7f64148962a8df61fa7322c.jpeg?v=1) | `ev:deepgram·cross·input-handling` | | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/84e7dcb47c124944ad4547c508cf5ee5.png?v=1) | `ev:deepgram·overlapping-meeting-speech-with-cross-talk·input-handling` | | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/353dc535361142fbab8920650fb7d1ca.png?v=1) | `ev:deepgram·bilingual-spanish-english-code-switching-speech·input-handling` | | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3f46ad25af044fcea47435b057d24056.png?v=1) | `ev:deepgram·medical-anatomy-narration-with-dense-jargon·input-handling` | | Output quality | Overlapping meeting speech with cross-talk | ⚠ struggled | 36.3/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1) | `ev:deepgram·overlapping-meeting-speech-with-cross-talk·output-quality` | | Output quality | Bilingual Spanish-English code-switching speech | ✗ failed | 38.1/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1) | `ev:deepgram·bilingual-spanish-english-code-switching-speech·output-quality` | | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 5.4/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1) | `ev:deepgram·medical-anatomy-narration-with-dense-jargon·output-quality` | | Output quality | cross-scenario | ◐ mixed | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-deepgram-65b5b534484c.md) | `ev:deepgram·cross·output-quality` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. > **Good batch STT metadata, uneven hard-case accuracy** > > Deepgram Nova-3 kept a stable developer payload and was excellent on the medical-jargon clip, but crosstalk and bilingual code-switching both produced high WER and lots of insertions and deletions. As configured here, it looks like a solid batch transcription API for mostly monolingual technical audio, not a safe default for overlapping meetings or Spanish-English conversation. ## Demo Recording [Video: Deepgram demo recording (download MP4)](https://cdn.futuresmart.ai/public/aidemos/22bd0b0611b646319c4ef9ba2fa9122f.mov?v=1) [▶️ Watch (streaming)](https://stream.futuresmart.ai/embed/38cd40e7-094a-43a0-b1c8-f5e51547d4ba) *Video — Tutorial recording from the benchmark task* ## Feature-by-Feature Breakdown ### Batch Audio Transcription **Verdict:** Mixed: strong on jargon, weak on overlap and code-switching. Converts uploaded pre-recorded audio into text in a single batch request. The member cards exercised it on overlapping meetings, medical-jargon narration, Spanish-English code-switching, and other prerecorded clips. **Input:** > **Audio** **Output:** > **File** **Input:** > **Audio** **Output:** > **File** **Input:** > **Audio** **Output:** > **File** **Bottom line:** Good on the medical narration clip, but not dependable on crosstalk or Spanish-English code-switching as configured here. ### Structured Transcript Output **Verdict:** Consistent developer payload across all three runs. Returns transcript results as structured JSON with fields like word-level timestamps, confidence values, punctuation, and speaker labels. The member cards exercised this across multiple runs and options such as smart_format, diarize, punctuate, and utterance. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** This is the most consistent part of the product: the JSON shape stayed stable across all three inputs. ### Speaker Diarization **Verdict:** Useful for speaker counts, but attribution correctness is unproven. Detects multiple speakers and includes speaker labels in the transcript payload. It was exercised on crosstalk and bilingual audio, where the responses surfaced multiple speaker labels. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Useful for speaker-count metadata, but the benchmark does not prove speaker-attribution accuracy. ## Reported vendor tiers The report notes that the served pricing page did not fully confirm the pre-recorded Nova-3 tab, so treat model-specific rates as unverified. | Plan | Price | Notes | | --- | --- | --- | | Pay As You Go | $200 free credit, then pay-as-you-go | No minimums, no expiration, no credit card required. STT concurrency up to 50 REST, 150 WSS, 5 Whisper Cloud. | | Growth | $4,000+ / year prepaid credits | Save up to 20%; 10% overage fee. STT concurrency up to 50 REST, 225 WSS, 5 Whisper Cloud. | | Enterprise | Requires sales contact | Large volume, data/deployment requirements, support; custom models, self-hosted/VPC, SLAs. | *Benchmark cost was calculated from list price × measured duration; the task header also reports $0.0063/min ($0.378/audio-hour) for the tested configuration.* ## Is It Right For You? **Use it if** - you need a batch STT API that returns word-level timing, confidence, and speaker labels - you mainly transcribe single-language technical narration and care about jargon recall - you care more about batch throughput than live streaming **Skip it if** - you need reliable overlapping-speech handling for crosstalk-heavy meetings - you need strong Spanish-English code-switching or broader multilingual transcription - you need validated diarization attribution rather than just speaker counts - you need measured streaming latency for a live benchmark ## Classification - **Category:** audio-speech - **Subcategory:** audio-to-text - **Type:** text - **Built for:** Other ## Frequently Asked Questions **Q: How did Deepgram accept the audio in this benchmark?** The runs used a single raw-body POST with audio in the request body and query-string options such as model, smart_format, diarize, punctuate, and utterances. The report says there was no form wrapper, and the vendor docs describe this as a single-call protocol. **Q: Does Deepgram return word timestamps, confidence, and speaker labels?** Yes. The raw response previews show word_timestamps, confidence, and speaker_labels detected on all three inputs, and the payload depth stayed at 3/3 each time. **Q: How accurate was Deepgram on overlapping speech?** On the crosstalk input, it scored 36.27% WER, returned 6,946 words against a 7,579-word reference, and produced 510 insertions and 1,143 deletions. It did return 4 speaker labels, but the report does not prove speaker-attribution correctness. **Q: How did Deepgram handle medical jargon?** It did well on the medical-jargon clip: 5.43% WER, 2,726 words returned against 2,728 reference words, and 100.0% jargon recall on the scored terms. **Q: How did Deepgram handle Spanish-English code-switching?** Poorly. On the bilingual clip it scored 38.13% WER, returned 5,691 words against 6,517 reference words, and Spanish token recall was only 3.8% (3 of 80 types). **Q: What did Deepgram cost and how fast was it?** The benchmark reported $0.22498 on crosstalk, $0.11802 on the medical clip, and $0.20354 on the bilingual clip. Latency ranged from 12.02s to 63.87s, with RTF from 0.01069 to 0.02981; streaming latency was not measured because this was a batch benchmark. ## Similar Tools AI tools similar to Deepgram: - [Gladia](https://aidemos.com/tools/gladia) — Batch STT with rich word-level metadata and strong jargon recall, but overlap handling is weak and bilingual WER needs a mono-downmixed rerun. - [AssemblyAI](https://aidemos.com/tools/assemblyai-speech-to-text) — Fast batch speech-to-text with rich metadata, strong jargon and mixed-language results, but overlap-heavy meetings can still lose too much. - [Speechmatics](https://aidemos.com/tools/speechmatics) — Strong batch STT for hard English audio, but weak on code-switching as configured. - [OpenAI](https://aidemos.com/tools/openai) — Batch speech-to-text with word timestamps, but a strict upload cap and weak multilingual performance make it a mixed fit for hard audio. - [Google Cloud Speech-to-Text](https://aidemos.com/tools/google-cloud-speech-to-text) — Timed batch transcripts for mostly English, jargon-heavy audio — but not for diarization or code-switching. - [AWS Transcribe](https://aidemos.com/tools/aws-transcribe) — Batch speech-to-text with word-level metadata, but accuracy drops on overlap and code-switching. - [ElevenLabs](https://aidemos.com/tools/elevenlabs) — Natural-sounding voice cloning and narration, but with only approximate voice identity. - [Rev AI](https://aidemos.com/tools/rev-ai) — Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on crosstalk and code-switching. - [ElevenLabs Scribe](https://aidemos.com/tools/elevenlabs-scribe) — Fast batch speech-to-text with word-level metadata, strongest on jargon and weaker on overlap/code-switching. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text, audio transcription, or speaker diarization system for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).