Deepgram  icon
audio-speech

Deepgram

Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching.

Visit Deepgram
Batch APIWord timestampsSpeaker labelsHard-audio benchmark
TL;DR — our verdictUpdated August 2026 · 9 test artifacts

Strong on technical narration, but not reliable on bilingual or overlapping speech.

Where it wins
  • you need a batch STT API that returns word-level timing, confidence, and speaker labels
  • you primarily transcribe single-language technical narration and care about jargon recall
  • you can process audio offline and care more about throughput than live streaming
Main limitation
  • you need strong Spanish-English code-switching or multilingual transcription
Pricing (verified plans)
Pay As You Go $200 free credit, then pay-as-you-goGrowth $4,000+ / year prepaid creditsEnterprise Requires sales contact
Strongest test artifacts

Our take

Deepgram Nova-3 returned a full developer payload on every run and was very strong on the medical-jargon clip, but it struggled badly on overlapping speech and bilingual code-switching. As configured here, it looks best for single-language technical audio, not multilingual or crosstalk-heavy recordings.

Tutorial recording of the benchmark workflow.

In-Depth Review

Our detailed analysis of Deepgram — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Audio Transcription
Handled the file, but crosstalk quality was poor.
Test Summary
Feature tested: Audio Transcription
Result: Failed — Handled the file, but crosstalk quality was poor.

Feature tested: Audio Transcription

Result: Failed

Verdict: Handled the file, but crosstalk quality was poor.

Expected behavior: Deepgram transcribes spoken audio into text from pre-recorded uploads and other difficult inputs. The member cards exercise it on a four-speaker crosstalk meeting, a medical narration with jargon, and a spontaneous Spanish-English conversation, plus batch-uploaded clips.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a). — crosstalk.wav

Observed output: Output artifact (Image): On the overlapping-speech clip, the run produced a transcript but the benchmark measured 36.27% WER with 1,143 deletions and 510 insertions, so the result was weak on crosstalk. — 04-transcript-detail.png

Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a). — crosstalk.wav

Output artifact: Output artifact (Image): On the overlapping-speech clip, the run produced a transcript but the benchmark measured 36.27% WER with 1,143 deletions and 510 insertions, so the result was weak on crosstalk. — 04-transcript-detail.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Medical Jargon test clip: medical_terms.mp3 (18:44, 8.58 MB, Gray's Anatomy via LibriVox). — medical_terms.mp3

Observed output: Output artifact (Image): On the medical-jargon clip, the run was much stronger, with 5.43% WER and 100% recall on the scored jargon terms. — 04-transcript-detail-2.png

Input artifact: Input artifact (Audio file): Medical Jargon test clip: medical_terms.mp3 (18:44, 8.58 MB, Gray's Anatomy via LibriVox). — medical_terms.mp3

Output artifact: Output artifact (Image): On the medical-jargon clip, the run was much stronger, with 5.43% WER and 100% recall on the scored jargon terms. — 04-transcript-detail-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Bilingual Code-Switching test clip: mix_language.mp3 (32:18, 22.19 MB, Bangor Miami herring1). — mix_language.mp3

Observed output: Output artifact (Image): On the bilingual clip, the run struggled badly, with 38.13% WER and only 3.8% Spanish token recall. — 04-transcript-detail-3.png

Input artifact: Input artifact (Audio file): Bilingual Code-Switching test clip: mix_language.mp3 (32:18, 22.19 MB, Bangor Miami herring1). — mix_language.mp3

Output artifact: Output artifact (Image): On the bilingual clip, the run struggled badly, with 38.13% WER and only 3.8% Spanish token recall. — 04-transcript-detail-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Functionally completed the transcription, but the overlap error rate was too high for reliable captions or meeting notes.

Deepgram transcribes spoken audio into text from pre-recorded uploads and other difficult inputs. The member cards exercise it on a four-speaker crosstalk meeting, a medical narration with jargon, and a spontaneous Spanish-English conversation, plus batch-uploaded clips.

audio
0:00 / 0:00
Loading audio...
Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a).
image
Output artifact for "Audio Transcription" test: On the overlapping-speech clip, the run produced a transcript but the benchmark measured 36.27% WER with 1,143 deletions and 510 insertions, so the result was weak on crosstalk., 04-transcript-detail.png
On the overlapping-speech clip, the run produced a transcript but the benchmark measured 36.27% WER with 1,143 deletions and 510 insertions, so the result was weak on crosstalk.
audio
0:00 / 0:00
Loading audio...
Medical Jargon test clip: medical_terms.mp3 (18:44, 8.58 MB, Gray's Anatomy via LibriVox).
image
Output artifact for "Audio Transcription" test: On the medical-jargon clip, the run was much stronger, with 5.43% WER and 100% recall on the scored jargon terms., 04-transcript-detail-2.png
On the medical-jargon clip, the run was much stronger, with 5.43% WER and 100% recall on the scored jargon terms.
audio
0:00 / 0:00
Loading audio...
Bilingual Code-Switching test clip: mix_language.mp3 (32:18, 22.19 MB, Bangor Miami herring1).
image
Output artifact for "Audio Transcription" test: On the bilingual clip, the run struggled badly, with 38.13% WER and only 3.8% Spanish token recall., 04-transcript-detail-3.png
On the bilingual clip, the run struggled badly, with 38.13% WER and only 3.8% Spanish token recall.
Bottom Line
Functionally completed the transcription, but the overlap error rate was too high for reliable captions or meeting notes.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Structured Transcript Output
Structured JSON transcript output returned on every run.
Test Summary
Feature tested: Structured Transcript Output
Result: Passed — Structured JSON transcript output returned on every run.

Feature tested: Structured Transcript Output

Result: Passed

Verdict: Structured JSON transcript output returned on every run.

Expected behavior: Deepgram can return a structured transcript payload from audio when options like smart_format, diarize, punctuate, and utterances are enabled. The exercised outputs included word-level timing, confidence values, and speaker labels.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT 1 — Overlapping Speech / Crosstalk; crosstalk.wav; 65.39 MB; 2142.709 s; AMI EN2002a. — crosstalk.wav

Observed output: Output artifact (Image): Raw API response preview for the overlap clip; the response surface exposed developer features such as word_timestamps, confidence, and speaker_labels. — 02-response-raw.png

Input artifact: Input artifact (Audio file): INPUT 1 — Overlapping Speech / Crosstalk; crosstalk.wav; 65.39 MB; 2142.709 s; AMI EN2002a. — crosstalk.wav

Output artifact: Output artifact (Image): Raw API response preview for the overlap clip; the response surface exposed developer features such as word_timestamps, confidence, and speaker_labels. — 02-response-raw.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT 2 — Medical Jargon; medical_terms.mp3; 8.58 MB; 1123.971 s; Gray's Anatomy via LibriVox. — medical_terms.mp3

Observed output: Output artifact (Image): Raw API response preview for the medical clip; the response surface again exposed word_timestamps, confidence, and speaker_labels, with one distinct speaker. — 02-response-raw-2.png

Input artifact: Input artifact (Audio file): INPUT 2 — Medical Jargon; medical_terms.mp3; 8.58 MB; 1123.971 s; Gray's Anatomy via LibriVox. — medical_terms.mp3

Output artifact: Output artifact (Image): Raw API response preview for the medical clip; the response surface again exposed word_timestamps, confidence, and speaker_labels, with one distinct speaker. — 02-response-raw-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT 3 — Bilingual Code-Switching; mix_language.mp3; 22.19 MB; 1938.495 s; Bangor Miami herring1. — mix_language.mp3

Observed output: Output artifact (Image): Raw API response preview for the bilingual clip; the response surface again exposed the structured transcript payload and developer metadata. — 02-response-raw-3.png

Input artifact: Input artifact (Audio file): INPUT 3 — Bilingual Code-Switching; mix_language.mp3; 22.19 MB; 1938.495 s; Bangor Miami herring1. — mix_language.mp3

Output artifact: Output artifact (Image): Raw API response preview for the bilingual clip; the response surface again exposed the structured transcript payload and developer metadata. — 02-response-raw-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Consistently useful for downstream parsing, even when the transcript quality itself varies.

Deepgram can return a structured transcript payload from audio when options like smart_format, diarize, punctuate, and utterances are enabled. The exercised outputs included word-level timing, confidence values, and speaker labels.

audio
0:00 / 0:00
Loading audio...
INPUT 1 — Overlapping Speech / Crosstalk; crosstalk.wav; 65.39 MB; 2142.709 s; AMI EN2002a.
OUTPUT
Output artifact for "Structured Transcript Output" test: Raw API response preview for the overlap clip; the response surface exposed developer features such as word_timestamps, confidence, and speaker_labels., 02-response-raw.png
Raw API response preview for the overlap clip; the response surface exposed developer features such as word_timestamps, confidence, and speaker_labels.
audio
0:00 / 0:00
Loading audio...
INPUT 2 — Medical Jargon; medical_terms.mp3; 8.58 MB; 1123.971 s; Gray's Anatomy via LibriVox.
OUTPUT
Output artifact for "Structured Transcript Output" test: Raw API response preview for the medical clip; the response surface again exposed word_timestamps, confidence, and speaker_labels, with one distinct speaker., 02-response-raw-2.png
Raw API response preview for the medical clip; the response surface again exposed word_timestamps, confidence, and speaker_labels, with one distinct speaker.
audio
0:00 / 0:00
Loading audio...
INPUT 3 — Bilingual Code-Switching; mix_language.mp3; 22.19 MB; 1938.495 s; Bangor Miami herring1.
OUTPUT
Output artifact for "Structured Transcript Output" test: Raw API response preview for the bilingual clip; the response surface again exposed the structured transcript payload and developer metadata., 02-response-raw-3.png
Raw API response preview for the bilingual clip; the response surface again exposed the structured transcript payload and developer metadata.
Bottom Line
Consistently useful for downstream parsing, even when the transcript quality itself varies.
From our researchearlier researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark
Speaker Diarization
Detected speaker labels and matched the overlap clip's speaker count.
Test Summary
Feature tested: Speaker Diarization
Result: Partial — Detected speaker labels and matched the overlap clip's speaker count.

Feature tested: Speaker Diarization

Result: Partial

Verdict: Detected speaker labels and matched the overlap clip's speaker count.

Expected behavior: Deepgram can label speakers in the transcript and report speaker counts when diarization is enabled. The exercised inputs included overlapping multi-speaker audio as well as single-speaker recordings.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a). — crosstalk.wav

Observed output: Output artifact (Image): The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference. — 03-terminal-metrics.png

Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a). — crosstalk.wav

Output artifact: Output artifact (Image): The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference. — 03-terminal-metrics.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Medical Jargon test clip: medical_terms.mp3 (18:44, 8.58 MB, Gray's Anatomy via LibriVox). — medical_terms.mp3

Observed output: Output artifact (Image): The raw response preview shows speaker_labels enabled and a single distinct speaker on the medical narration. — 02-response-raw-2.png

Input artifact: Input artifact (Audio file): Medical Jargon test clip: medical_terms.mp3 (18:44, 8.58 MB, Gray's Anatomy via LibriVox). — medical_terms.mp3

Output artifact: Output artifact (Image): The raw response preview shows speaker_labels enabled and a single distinct speaker on the medical narration. — 02-response-raw-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Bilingual Code-Switching test clip: mix_language.mp3 (32:18, 22.19 MB, Bangor Miami herring1). — mix_language.mp3

Observed output: Output artifact (Image): The raw response preview shows speaker_labels enabled on the bilingual clip and a successful structured transcript payload. — 02-response-raw-3.png

Input artifact: Input artifact (Audio file): Bilingual Code-Switching test clip: mix_language.mp3 (32:18, 22.19 MB, Bangor Miami herring1). — mix_language.mp3

Output artifact: Output artifact (Image): The raw response preview shows speaker_labels enabled on the bilingual clip and a successful structured transcript payload. — 02-response-raw-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Speaker labels are present and count-correct on the hardest overlap case, but attribution correctness was not independently scored.

Deepgram can label speakers in the transcript and report speaker counts when diarization is enabled. The exercised inputs included overlapping multi-speaker audio as well as single-speaker recordings.

audio
0:00 / 0:00
Loading audio...
Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a).
image
Output artifact for "Speaker Diarization" test: The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference., 03-terminal-metrics.png
The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference.
audio
0:00 / 0:00
Loading audio...
Medical Jargon test clip: medical_terms.mp3 (18:44, 8.58 MB, Gray's Anatomy via LibriVox).
image
Output artifact for "Speaker Diarization" test: The raw response preview shows speaker_labels enabled and a single distinct speaker on the medical narration., 02-response-raw-2.png
The raw response preview shows speaker_labels enabled and a single distinct speaker on the medical narration.
audio
0:00 / 0:00
Loading audio...
Bilingual Code-Switching test clip: mix_language.mp3 (32:18, 22.19 MB, Bangor Miami herring1).
image
Output artifact for "Speaker Diarization" test: The raw response preview shows speaker_labels enabled on the bilingual clip and a successful structured transcript payload., 02-response-raw-3.png
The raw response preview shows speaker_labels enabled on the bilingual clip and a successful structured transcript payload.
Bottom Line
Speaker labels are present and count-correct on the hardest overlap case, but attribution correctness was not independently scored.
From our researchearlier research

How it scored on the research's own criteria

The 4 evaluation dimensions from our hands-on research on Deepgram , each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Output qualityWeak2/5One narration run was very accurate, but the other two were far weaker, with high WER and large deletion/insertion counts; the Spanish-English mix was especially brittle, so the overall accuracy picture is poor rather than merely uneven.open proof ↗
Automation levelStrong5/5Each run went through in one shot with no handholding, using the same simple request pattern from start to finish, so the tool deserves the top score for automation.
ExportStrong5/5Every run came back with word-level timing, confidence, and speaker labels in a deep JSON payload, so the tool consistently exposes the full transcript package instead of a thin text-only result.
Input handlingStrong5/5It accepted every file without complaint and finished each run quickly, with all three wall-clock times well under real time and the listed costs reported, so this is a clear 5/5 for basic ingestion and throughput.open proof ↗

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Published Deepgram plans

Vendor pricing was available, but the pre-recorded Nova-3 tab was not fully browser-verified in this research.

Pay As You Go
$200 free credit, then pay-as-you-go
No minimums, no expiration, no credit card required; STT concurrency up to 50 REST, 150 WSS, 5 Whisper Cloud.
Growth
$4,000+ / year prepaid credits
Save up to 20%; 10% overage fee; STT concurrency up to 50 REST, 225 WSS, 5 Whisper Cloud.
Enterprise
Requires sales contact
Large volume, custom models, self-hosted/VPC, SLAs, and deployment requirements.

Benchmark cost here reflects the configured list price of $0.0063/min ($0.378/audio-hour), not an invoice. The vendor page also shows add-ons billed separately, including Speaker Diarization at $0.0020/min on Pay As You Go. Pre-recorded-specific Nova-3 pricing should be re-verified in a live browser before publishing.

✓ Use This If
you need a batch STT API that returns word-level timing, confidence, and speaker labels
you primarily transcribe single-language technical narration and care about jargon recall
you can process audio offline and care more about throughput than live streaming
✕ Skip This If
you need strong Spanish-English code-switching or multilingual transcription
you need audited diarization attribution quality, not just speaker counts
you need measured streaming latency for a live transcription benchmark
audio-speechaudio-to-textspeechOther
Yes. Across all three benchmark runs, the response previews showed word-level timing, confidence values, and speaker labels in the structured JSON payload.
On the four-speaker overlap clip, it scored 36.27% WER with 510 insertions and 1143 deletions. It detected four speaker labels, but the transcript quality was still poor because of crosstalk.
This was its best result in the set. On the medical narration clip it scored 5.43% WER and recalled all 9 scored jargon terms.
Poorly as configured. On the bilingual clip it scored 38.13% WER and Spanish token recall was only 3.8% (3 out of 80 types).
The benchmark used a configured list price of $0.0063/min, which equals $0.378 per audio-hour. The vendor pricing page also shows add-ons billed separately, and the pre-recorded-specific Nova-3 rate was not fully browser-verified.
No. This was a batch benchmark only, and the report explicitly says streaming latency was not measured.

Banner Preview

How the embed badge will look on your site

Deepgram  featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/deepgram-nova-3?utm_source=deepgram-nova-3_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="Deepgram | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like Deepgram to enhance your workflow.

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom speech-to-text transcription, audio transcription, or transcription workflow for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top