--- title: "Deepgram Nova-3" type: "AI Tool" url: "https://aidemos.com/tools/deepgram-nova-3" description: "We ran batch speech-to-text on medical, overlapping, and bilingual clips; Nova-3 nailed jargon and metadata, but broke on crosstalk and code-switching." category: "audio-speech" website: "https://deepgram.com" published: "2026-08-20T12:32:56.034069+00:00" updated: "2026-08-20T12:46:06.762580+00:00" --- # Deepgram Nova-3 Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching. ## TL;DR Verdict **Strong on technical narration, but not reliable on bilingual or overlapping speech.** **Where it wins:** - you need a batch STT API that returns word-level timing, confidence, and speaker labels - you primarily transcribe single-language technical narration and care about jargon recall - you can process audio offline and care more about throughput than live streaming **Main limitation:** you need strong Spanish-English code-switching or multilingual transcription **Pricing:** Pay As You Go $200 free credit, then pay-as-you-go · Growth $4,000+ / year prepaid credits · Enterprise Requires sales contact `Batch API` · `Word timestamps` · `Speaker labels` · `Hard-audio benchmark` **Website:** [Visit Deepgram Nova-3](https://deepgram.com) > **Strong on technical narration, but not reliable on bilingual or overlapping speech.** > > Deepgram Nova-3 returned a full developer payload on every run and was very strong on the medical-jargon clip, but it struggled badly on overlapping speech and bilingual code-switching. As configured here, it looks best for single-language technical audio, not multilingual or crosstalk-heavy recordings. ## Demo Recording [Video: Deepgram Nova-3 demo recording](https://cdn.futuresmart.ai/public/aidemos/e5c976315a46469eb36cd8cbacffe892.mov?v=1) *Video — Tutorial recording of the benchmark workflow.* ## Feature-by-Feature Breakdown ### Audio Transcription **Verdict:** Handled the file, but crosstalk quality was poor. Deepgram transcribes spoken audio into text from pre-recorded uploads and other difficult inputs. The member cards exercise it on a four-speaker crosstalk meeting, a medical narration with jargon, and a spontaneous Spanish-English conversation, plus batch-uploaded clips. **Input:** 1 > **Audio** — 1 **Output:** Transcript detail > **Image** — Transcript detail **Input:** 2 > **Audio** — 2 **Output:** Transcript detail > **Image** — Transcript detail **Input:** 3 > **Audio** — 3 **Output:** Transcript detail > **Image** — Transcript detail **Bottom line:** Functionally completed the transcription, but the overlap error rate was too high for reliable captions or meeting notes. ### Structured Transcript Output **Verdict:** Structured JSON transcript output returned on every run. Deepgram can return a structured transcript payload from audio when options like smart_format, diarize, punctuate, and utterances are enabled. The exercised outputs included word-level timing, confidence values, and speaker labels. **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Bottom line:** Consistently useful for downstream parsing, even when the transcript quality itself varies. ### Speaker Diarization **Verdict:** Detected speaker labels and matched the overlap clip's speaker count. Deepgram can label speakers in the transcript and report speaker counts when diarization is enabled. The exercised inputs included overlapping multi-speaker audio as well as single-speaker recordings. **Input:** 1 > **Audio** — 1 **Output:** Run metrics > **Image** — Run metrics **Input:** 2 > **Audio** — 2 **Output:** Raw API response > **Image** — Raw API response **Input:** 3 > **Audio** — 3 **Output:** Raw API response > **Image** — Raw API response **Bottom line:** Speaker labels are present and count-correct on the hardest overlap case, but attribution correctness was not independently scored. ## Published Deepgram plans Vendor pricing was available, but the pre-recorded Nova-3 tab was not fully browser-verified in this research. | Plan | Price | Notes | | --- | --- | --- | | Pay As You Go | $200 free credit, then pay-as-you-go | No minimums, no expiration, no credit card required; STT concurrency up to 50 REST, 150 WSS, 5 Whisper Cloud. | | Growth | $4,000+ / year prepaid credits | Save up to 20%; 10% overage fee; STT concurrency up to 50 REST, 225 WSS, 5 Whisper Cloud. | | Enterprise | Requires sales contact | Large volume, custom models, self-hosted/VPC, SLAs, and deployment requirements. | *Benchmark cost here reflects the configured list price of $0.0063/min ($0.378/audio-hour), not an invoice. The vendor page also shows add-ons billed separately, including Speaker Diarization at $0.0020/min on Pay As You Go. Pre-recorded-specific Nova-3 pricing should be re-verified in a live browser before publishing.* ## Is It Right For You? **Use it if** - you need a batch STT API that returns word-level timing, confidence, and speaker labels - you primarily transcribe single-language technical narration and care about jargon recall - you can process audio offline and care more about throughput than live streaming **Skip it if** - you need strong Spanish-English code-switching or multilingual transcription - you need audited diarization attribution quality, not just speaker counts - you need measured streaming latency for a live transcription benchmark ## Classification - **Category:** audio-speech - **Subcategory:** audio-to-text - **Type:** speech - **Built for:** Other ## Frequently Asked Questions **Q: Does Deepgram Nova-3 return speaker labels and word timestamps?** Yes. Across all three benchmark runs, the response previews showed word-level timing, confidence values, and speaker labels in the structured JSON payload. **Q: How accurate was it on overlapping speech?** On the four-speaker overlap clip, it scored 36.27% WER with 510 insertions and 1143 deletions. It detected four speaker labels, but the transcript quality was still poor because of crosstalk. **Q: How did it handle medical jargon?** This was its best result in the set. On the medical narration clip it scored 5.43% WER and recalled all 9 scored jargon terms. **Q: How did it handle Spanish-English code-switching?** Poorly as configured. On the bilingual clip it scored 38.13% WER and Spanish token recall was only 3.8% (3 out of 80 types). **Q: What did it cost per audio hour in this benchmark?** The benchmark used a configured list price of $0.0063/min, which equals $0.378 per audio-hour. The vendor pricing page also shows add-ons billed separately, and the pre-recorded-specific Nova-3 rate was not fully browser-verified. **Q: Was streaming latency measured?** No. This was a batch benchmark only, and the report explicitly says streaming latency was not measured. ## Similar Tools AI tools similar to Deepgram Nova-3: - [Rev AI](https://aidemos.com/tools/rev-ai) — Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on hard audio and weak multilingual recall. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text transcription, audio transcription, or transcription workflow for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).