Deepgram  icon
audio-speech

Deepgram

Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching.

Visit Deepgram
Batch STTWord timestampsSpeaker labelsWeak code-switching
TL;DR — our verdictUpdated September 2026 · 13 test artifacts

Mixed results, strong metadata

Where it wins
  • you need a batch STT API that returns word-level timing, confidence, and speaker labels
  • you primarily transcribe single-language technical narration and care about jargon recall
  • you can process audio offline and care more about throughput than live streaming
Main limitation
  • you need strong Spanish-English code-switching or multilingual transcription
Pricing (verified plans)
Pay As You Go $200 free credit, then pay-as-you-goGrowth $4,000+ / year prepaid creditsEnterprise Requires sales contact
Strongest test artifacts

Our take

Deepgram Nova-3 returned a full developer payload on every run and was excellent on the medical-jargon clip, but the crosstalk and bilingual runs were noisy enough that I would not treat it as reliable for overlapping or code-switched audio as configured here.

Screen recording of the benchmark workflow and run review for Deepgram Nova-3.

In-Depth Review

Our detailed analysis of Deepgram — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Audio Transcription
Mixed
Test Summary
Feature tested: Audio Transcription
Result: Failed — Mixed

Feature tested: Audio Transcription

Result: Failed

Verdict: Mixed

Expected behavior: Deepgram converts uploaded pre-recorded spoken audio into text through a single API request. The member cards exercised it on benchmark clips including overlapping meetings, medical-jargon narration, Spanish-English code-switching, and other prerecorded files.

Test case: Image → Text/code file

Input type: Image

Input used: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, 65.39 MB, 2142.709 s, four-way overlapping meeting audio. — 84e7dcb47c124944ad4547c508cf5ee5.png

Observed output: Output artifact (Text/code file): Crosstalk run: Deepgram returned a structured JSON transcript for crosstalk.wav; the scored run posted 36.27% WER, 63.87s latency, 0.02981 RTF, 6,946 returned words, and 4 detected speakers. — raw-response.json

Input artifact: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, 65.39 MB, 2142.709 s, four-way overlapping meeting audio. — 84e7dcb47c124944ad4547c508cf5ee5.png

Output artifact: Output artifact (Text/code file): Crosstalk run: Deepgram returned a structured JSON transcript for crosstalk.wav; the scored run posted 36.27% WER, 63.87s latency, 0.02981 RTF, 6,946 returned words, and 4 detected speakers. — raw-response.json

What changed: Image transformed into Text/code file

Test case: Image → Text/code file

Input type: Image

Input used: Input artifact (Image): Medical Jargon — medical_terms.mp3, 18:44, 8.58 MB, 1123.971 s, single-speaker narration dense with anatomical terms. — 3f46ad25af044fcea47435b057d24056.png

Observed output: Output artifact (Text/code file): Medical-jargon run: Deepgram returned a transcript JSON for medical_terms.mp3; the scored run posted 5.43% WER, 12.02s latency, 0.01069 RTF, 2,726 returned words against 2,728 reference words, and 100% jargon recall. — raw-response-2.json

Input artifact: Input artifact (Image): Medical Jargon — medical_terms.mp3, 18:44, 8.58 MB, 1123.971 s, single-speaker narration dense with anatomical terms. — 3f46ad25af044fcea47435b057d24056.png

Output artifact: Output artifact (Text/code file): Medical-jargon run: Deepgram returned a transcript JSON for medical_terms.mp3; the scored run posted 5.43% WER, 12.02s latency, 0.01069 RTF, 2,726 returned words against 2,728 reference words, and 100% jargon recall. — raw-response-2.json

What changed: Image transformed into Text/code file

Test case: Image → Text/code file

Input type: Image

Input used: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, 32:18, 22.19 MB, 1938.495 s, spontaneous Spanish-English conversation. — 353dc535361142fbab8920650fb7d1ca.png

Observed output: Output artifact (Text/code file): Bilingual run: Deepgram returned a transcript JSON for mix_language.mp3; the scored run posted 38.13% WER, 27.58s latency, 0.01423 RTF, 5,691 returned words against 6,517 reference words, and Spanish recall was only 3.8%. — raw-response-3.json

Input artifact: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, 32:18, 22.19 MB, 1938.495 s, spontaneous Spanish-English conversation. — 353dc535361142fbab8920650fb7d1ca.png

Output artifact: Output artifact (Text/code file): Bilingual run: Deepgram returned a transcript JSON for mix_language.mp3; the scored run posted 38.13% WER, 27.58s latency, 0.01423 RTF, 5,691 returned words against 6,517 reference words, and Spanish recall was only 3.8%. — raw-response-3.json

What changed: Image transformed into Text/code file

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, batch throughput and latency run. — 84e7dcb47c124944ad4547c508cf5ee5.png

Observed output: Output artifact (Image): Run metrics for crosstalk.wav show 36.27% WER, 63.87s wall-clock latency, 0.02981 RTF, and an estimated cost of $0.22498 for the batch run. — 03-terminal-metrics-input-1.png

Input artifact: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, batch throughput and latency run. — 84e7dcb47c124944ad4547c508cf5ee5.png

Output artifact: Output artifact (Image): Run metrics for crosstalk.wav show 36.27% WER, 63.87s wall-clock latency, 0.02981 RTF, and an estimated cost of $0.22498 for the batch run. — 03-terminal-metrics-input-1.png

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Medical Jargon — medical_terms.mp3, batch throughput and latency run. — 3f46ad25af044fcea47435b057d24056.png

Observed output: Output artifact (Image): Run metrics for medical_terms.mp3 show 5.43% WER, 12.02s wall-clock latency, 0.01069 RTF, and an estimated cost of $0.11802 for the batch run. — 03-terminal-metrics-input-2.png

Input artifact: Input artifact (Image): Medical Jargon — medical_terms.mp3, batch throughput and latency run. — 3f46ad25af044fcea47435b057d24056.png

Output artifact: Output artifact (Image): Run metrics for medical_terms.mp3 show 5.43% WER, 12.02s wall-clock latency, 0.01069 RTF, and an estimated cost of $0.11802 for the batch run. — 03-terminal-metrics-input-2.png

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, batch throughput and latency run. — 353dc535361142fbab8920650fb7d1ca.png

Observed output: Output artifact (Image): Run metrics for mix_language.mp3 show 38.13% WER, 27.58s wall-clock latency, 0.01423 RTF, and an estimated cost of $0.20354 for the batch run. — 03-terminal-metrics-input-3.png

Input artifact: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, batch throughput and latency run. — 353dc535361142fbab8920650fb7d1ca.png

Output artifact: Output artifact (Image): Run metrics for mix_language.mp3 show 38.13% WER, 27.58s wall-clock latency, 0.01423 RTF, and an estimated cost of $0.20354 for the batch run. — 03-terminal-metrics-input-3.png

What changed: Image transformed into Image

Why it matters / Conclusion: Useful for batch transcription when the audio is mostly single-speaker and monolingual, but not reliable for crosstalk-heavy meetings or code-switched conversation.

Deepgram converts uploaded pre-recorded spoken audio into text through a single API request. The member cards exercised it on benchmark clips including overlapping meetings, medical-jargon narration, Spanish-English code-switching, and other prerecorded files.

image
Input artifact for "Audio Transcription" test: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, 65.39 MB, 2142.709 s, four-way overlapping meeting audio., 84e7dcb47c124944ad4547c508cf5ee5.png
Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, 65.39 MB, 2142.709 s, four-way overlapping meeting audio.
json
raw-response.json
Loading file...
Crosstalk run: Deepgram returned a structured JSON transcript for crosstalk.wav; the scored run posted 36.27% WER, 63.87s latency, 0.02981 RTF, 6,946 returned words, and 4 detected speakers.
image
Input artifact for "Audio Transcription" test: Medical Jargon — medical_terms.mp3, 18:44, 8.58 MB, 1123.971 s, single-speaker narration dense with anatomical terms., 3f46ad25af044fcea47435b057d24056.png
Medical Jargon — medical_terms.mp3, 18:44, 8.58 MB, 1123.971 s, single-speaker narration dense with anatomical terms.
json
raw-response-2.json
Loading file...
Medical-jargon run: Deepgram returned a transcript JSON for medical_terms.mp3; the scored run posted 5.43% WER, 12.02s latency, 0.01069 RTF, 2,726 returned words against 2,728 reference words, and 100% jargon recall.
image
Input artifact for "Audio Transcription" test: Bilingual Code-Switching — mix_language.mp3, 32:18, 22.19 MB, 1938.495 s, spontaneous Spanish-English conversation., 353dc535361142fbab8920650fb7d1ca.png
Bilingual Code-Switching — mix_language.mp3, 32:18, 22.19 MB, 1938.495 s, spontaneous Spanish-English conversation.
json
raw-response-3.json
Loading file...
Bilingual run: Deepgram returned a transcript JSON for mix_language.mp3; the scored run posted 38.13% WER, 27.58s latency, 0.01423 RTF, 5,691 returned words against 6,517 reference words, and Spanish recall was only 3.8%.
image
Input artifact for "Audio Transcription" test: Overlapping Speech / Crosstalk — crosstalk.wav, batch throughput and latency run., 84e7dcb47c124944ad4547c508cf5ee5.png
Overlapping Speech / Crosstalk — crosstalk.wav, batch throughput and latency run.
image
Output artifact for "Audio Transcription" test: Run metrics for crosstalk.wav show 36.27% WER, 63.87s wall-clock latency, 0.02981 RTF, and an estimated cost of $0.22498 for the batch run., 03-terminal-metrics-input-1.png
Run metrics for crosstalk.wav show 36.27% WER, 63.87s wall-clock latency, 0.02981 RTF, and an estimated cost of $0.22498 for the batch run.
image
Input artifact for "Audio Transcription" test: Medical Jargon — medical_terms.mp3, batch throughput and latency run., 3f46ad25af044fcea47435b057d24056.png
Medical Jargon — medical_terms.mp3, batch throughput and latency run.
image
Output artifact for "Audio Transcription" test: Run metrics for medical_terms.mp3 show 5.43% WER, 12.02s wall-clock latency, 0.01069 RTF, and an estimated cost of $0.11802 for the batch run., 03-terminal-metrics-input-2.png
Run metrics for medical_terms.mp3 show 5.43% WER, 12.02s wall-clock latency, 0.01069 RTF, and an estimated cost of $0.11802 for the batch run.
image
Input artifact for "Audio Transcription" test: Bilingual Code-Switching — mix_language.mp3, batch throughput and latency run., 353dc535361142fbab8920650fb7d1ca.png
Bilingual Code-Switching — mix_language.mp3, batch throughput and latency run.
image
Output artifact for "Audio Transcription" test: Run metrics for mix_language.mp3 show 38.13% WER, 27.58s wall-clock latency, 0.01423 RTF, and an estimated cost of $0.20354 for the batch run., 03-terminal-metrics-input-3.png
Run metrics for mix_language.mp3 show 38.13% WER, 27.58s wall-clock latency, 0.01423 RTF, and an estimated cost of $0.20354 for the batch run.
Bottom Line
Useful for batch transcription when the audio is mostly single-speaker and monolingual, but not reliable for crosstalk-heavy meetings or code-switched conversation.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Structured Transcript Output
Consistent
Test Summary
Feature tested: Structured Transcript Output
Result: Passed — Consistent

Feature tested: Structured Transcript Output

Result: Passed

Verdict: Consistent

Expected behavior: Deepgram returns transcript payloads as structured JSON with fields such as word-level timing, confidence values, speaker labels, punctuated tokens, and other parseable metadata. The member cards exercised this output shape across runs and options like smart_format, diarize, punctuate, and utterances.

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, batch transcription run used to inspect JSON structure and metadata. — 84e7dcb47c124944ad4547c508cf5ee5.png

Observed output: Output artifact (Image): Raw API response excerpt for crosstalk.wav showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers. — 02-response-raw-input-1.png

Input artifact: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, batch transcription run used to inspect JSON structure and metadata. — 84e7dcb47c124944ad4547c508cf5ee5.png

Output artifact: Output artifact (Image): Raw API response excerpt for crosstalk.wav showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers. — 02-response-raw-input-1.png

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Medical Jargon — medical_terms.mp3, batch transcription run used to inspect JSON structure and metadata. — 3f46ad25af044fcea47435b057d24056.png

Observed output: Output artifact (Image): Raw API response excerpt for medical_terms.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 5,845 timed tokens, and 1 distinct speaker. — 02-response-raw-input-2.png

Input artifact: Input artifact (Image): Medical Jargon — medical_terms.mp3, batch transcription run used to inspect JSON structure and metadata. — 3f46ad25af044fcea47435b057d24056.png

Output artifact: Output artifact (Image): Raw API response excerpt for medical_terms.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 5,845 timed tokens, and 1 distinct speaker. — 02-response-raw-input-2.png

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, batch transcription run used to inspect JSON structure and metadata. — 353dc535361142fbab8920650fb7d1ca.png

Observed output: Output artifact (Image): Raw API response excerpt for mix_language.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 11,684 timed tokens, and 3 distinct speakers. — 02-response-raw-input-3.png

Input artifact: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, batch transcription run used to inspect JSON structure and metadata. — 353dc535361142fbab8920650fb7d1ca.png

Output artifact: Output artifact (Image): Raw API response excerpt for mix_language.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 11,684 timed tokens, and 3 distinct speakers. — 02-response-raw-input-3.png

What changed: Image transformed into Image

Why it matters / Conclusion: This is the most dependable part of the product in this benchmark: the output shape stayed consistent enough for downstream parsing on every clip.

Deepgram returns transcript payloads as structured JSON with fields such as word-level timing, confidence values, speaker labels, punctuated tokens, and other parseable metadata. The member cards exercised this output shape across runs and options like smart_format, diarize, punctuate, and utterances.

image
Input artifact for "Structured Transcript Output" test: Overlapping Speech / Crosstalk — crosstalk.wav, batch transcription run used to inspect JSON structure and metadata., 84e7dcb47c124944ad4547c508cf5ee5.png
Overlapping Speech / Crosstalk — crosstalk.wav, batch transcription run used to inspect JSON structure and metadata.
image
Output artifact for "Structured Transcript Output" test: Raw API response excerpt for crosstalk.wav showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers., 02-response-raw-input-1.png
Raw API response excerpt for crosstalk.wav showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers.
image
Input artifact for "Structured Transcript Output" test: Medical Jargon — medical_terms.mp3, batch transcription run used to inspect JSON structure and metadata., 3f46ad25af044fcea47435b057d24056.png
Medical Jargon — medical_terms.mp3, batch transcription run used to inspect JSON structure and metadata.
image
Output artifact for "Structured Transcript Output" test: Raw API response excerpt for medical_terms.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 5,845 timed tokens, and 1 distinct speaker., 02-response-raw-input-2.png
Raw API response excerpt for medical_terms.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 5,845 timed tokens, and 1 distinct speaker.
image
Input artifact for "Structured Transcript Output" test: Bilingual Code-Switching — mix_language.mp3, batch transcription run used to inspect JSON structure and metadata., 353dc535361142fbab8920650fb7d1ca.png
Bilingual Code-Switching — mix_language.mp3, batch transcription run used to inspect JSON structure and metadata.
image
Output artifact for "Structured Transcript Output" test: Raw API response excerpt for mix_language.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 11,684 timed tokens, and 3 distinct speakers., 02-response-raw-input-3.png
Raw API response excerpt for mix_language.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 11,684 timed tokens, and 3 distinct speakers.
Bottom Line
This is the most dependable part of the product in this benchmark: the output shape stayed consistent enough for downstream parsing on every clip.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Speaker Diarization
Partial
Test Summary
Feature tested: Speaker Diarization
Result: Partial — Partial

Feature tested: Speaker Diarization

Result: Partial

Verdict: Partial

Expected behavior: Deepgram can surface speaker labels in its transcript output and, on the hardest crosstalk clip, matched the four-participant count. The benchmark cards note that speaker attribution quality itself was not fully verified.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a). — crosstalk.wav

Observed output: Output artifact (Image): The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference. — 03-terminal-metrics.png

Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a). — crosstalk.wav

Output artifact: Output artifact (Image): The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference. — 03-terminal-metrics.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk — crosstalk.wav, four-way overlapping meeting audio with diarization stress. — crosstalk.wav

Observed output: Output artifact (Image): Transcript detail for the crosstalk clip shows the largest divergence, but also confirms diarization detection with four labels and four true participants; the transcript itself was still noisy, with a 36.27% WER and 510 insertions. — 04-transcript-detail-input-1.png

Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk — crosstalk.wav, four-way overlapping meeting audio with diarization stress. — crosstalk.wav

Output artifact: Output artifact (Image): Transcript detail for the crosstalk clip shows the largest divergence, but also confirms diarization detection with four labels and four true participants; the transcript itself was still noisy, with a 36.27% WER and 510 insertions. — 04-transcript-detail-input-1.png

What changed: Audio file transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Medical Jargon — medical_terms.mp3, single-speaker narration used to verify speaker-label detection. — 3f46ad25af044fcea47435b057d24056.png

Observed output: Output artifact (Image): Raw response excerpt for the medical-jargon clip shows speaker_labels yes and a single distinct speaker, which is consistent with the single-speaker source audio. — 02-response-raw-input-2.png

Input artifact: Input artifact (Image): Medical Jargon — medical_terms.mp3, single-speaker narration used to verify speaker-label detection. — 3f46ad25af044fcea47435b057d24056.png

Output artifact: Output artifact (Image): Raw response excerpt for the medical-jargon clip shows speaker_labels yes and a single distinct speaker, which is consistent with the single-speaker source audio. — 02-response-raw-input-2.png

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation used to verify speaker-label detection. — 353dc535361142fbab8920650fb7d1ca.png

Observed output: Output artifact (Image): Raw response excerpt for the bilingual clip shows speaker_labels yes and three distinct speakers, but the benchmark did not measure whether those labels were attributed to the right voices. — 02-response-raw-input-3.png

Input artifact: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation used to verify speaker-label detection. — 353dc535361142fbab8920650fb7d1ca.png

Output artifact: Output artifact (Image): Raw response excerpt for the bilingual clip shows speaker_labels yes and three distinct speakers, but the benchmark did not measure whether those labels were attributed to the right voices. — 02-response-raw-input-3.png

What changed: Image transformed into Image

Why it matters / Conclusion: Good for speaker-count metadata, but attribution quality remains unproven.

Deepgram can surface speaker labels in its transcript output and, on the hardest crosstalk clip, matched the four-participant count. The benchmark cards note that speaker attribution quality itself was not fully verified.

audio
0:00 / 0:00
Loading audio...
Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a).
image
Output artifact for "Speaker Diarization" test: The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference., 03-terminal-metrics.png
The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference.
audio
0:00 / 0:00
Loading audio...
Overlapping Speech / Crosstalk — crosstalk.wav, four-way overlapping meeting audio with diarization stress.
image
Output artifact for "Speaker Diarization" test: Transcript detail for the crosstalk clip shows the largest divergence, but also confirms diarization detection with four labels and four true participants; the transcript itself was still noisy, with a 36.27% WER and 510 insertions., 04-transcript-detail-input-1.png
Transcript detail for the crosstalk clip shows the largest divergence, but also confirms diarization detection with four labels and four true participants; the transcript itself was still noisy, with a 36.27% WER and 510 insertions.
image
Input artifact for "Speaker Diarization" test: Medical Jargon — medical_terms.mp3, single-speaker narration used to verify speaker-label detection., 3f46ad25af044fcea47435b057d24056.png
Medical Jargon — medical_terms.mp3, single-speaker narration used to verify speaker-label detection.
image
Output artifact for "Speaker Diarization" test: Raw response excerpt for the medical-jargon clip shows speaker_labels yes and a single distinct speaker, which is consistent with the single-speaker source audio., 02-response-raw-input-2.png
Raw response excerpt for the medical-jargon clip shows speaker_labels yes and a single distinct speaker, which is consistent with the single-speaker source audio.
image
Input artifact for "Speaker Diarization" test: Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation used to verify speaker-label detection., 353dc535361142fbab8920650fb7d1ca.png
Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation used to verify speaker-label detection.
image
Output artifact for "Speaker Diarization" test: Raw response excerpt for the bilingual clip shows speaker_labels yes and three distinct speakers, but the benchmark did not measure whether those labels were attributed to the right voices., 02-response-raw-input-3.png
Raw response excerpt for the bilingual clip shows speaker_labels yes and three distinct speakers, but the benchmark did not measure whether those labels were attributed to the right voices.
Bottom Line
Good for speaker-count metadata, but attribution quality remains unproven.
From our researchearlier researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark

How it scored on the research's own criteria

The 3 evaluation dimensions from our hands-on research on Deepgram , each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Output qualityWeak2/5The model is excellent on clean medical narration, but the other two runs are badly hurt by heavy error rates, and the bilingual case especially shows that it loses most Spanish even when the conversation is mostly English.open proof ↗
Automation levelStrong5/5Each run completed as a single automated request-response flow with no manual intervention, so the workflow is as hands-off as it can be in this setup.
Input handlingStrong5/5It handled all three audio files cleanly, stayed fast on each run, and kept costs modest, so this is consistently top-tier input handling rather than just passing the basics.open proof ↗

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Official pricing from the vendor page

The benchmark config was logged at $0.0063/min ($0.378/audio-hour); the report notes that the served pre-recorded Nova-3 tab could not be fully verified in-browser.

Pay As You Go
$200 free credit, then pay-as-you-go
No minimums, no expiration, no credit card required; STT concurrency up to 50 REST, 150 WSS, 5 Whisper Cloud.
Growth
$4,000+ / year prepaid credits
Save up to 20%; credits redeemed against actual usage; 10% overage fee; STT concurrency up to 50 REST, 225 WSS, 5 Whisper Cloud.
Enterprise
Requires sales contact
Large volume, custom models, self-hosted/VPC deployment, and SLAs are listed in the report.

The report also records Nova-3 Monolingual at $0.0048/min and Nova-3 Multilingual at $0.0058/min, with Speaker Diarization listed separately at $0.0020/min. Smart Formatting is included.

✓ Use This If
you need a batch STT API that returns word-level timing, confidence, and speaker labels
you primarily transcribe single-language technical narration and care about jargon recall
you can process audio offline and care more about throughput than live streaming
✕ Skip This If
you need strong Spanish-English code-switching or multilingual transcription
you need reliable overlap handling for crosstalk-heavy meetings
you need audited diarization attribution quality, not just speaker counts
you need measured streaming latency for a live transcription benchmark
audio-speechaudio-to-textspeechOther
Yes. All three benchmark runs reported word_timestamps, confidence, and speaker_labels, and the payload depth stayed at 3/3 each time.
On the crosstalk clip, it scored 36.27% WER with 510 insertions and 1,143 deletions. It did detect 4 speakers, but the transcript was noisy.
This was its best run. On the medical-jargon clip it scored 5.43% WER, returned 2,726 words against 2,728 in the reference, and achieved 100% jargon recall.
Poorly. On the bilingual clip it scored 38.13% WER and only 3.8% Spanish token recall, with 3 of 80 Spanish types recovered.
No. This benchmark was batch-only, and the report explicitly says streaming latency was not measured.
The benchmark configuration logs $0.0063/min, which is $0.378 per audio hour. The report also notes that the vendor's served pre-recorded Nova-3 pricing tab could not be fully verified in-browser.

Banner Preview

How the embed badge will look on your site

Deepgram  featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/deepgram?utm_source=deepgram_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="Deepgram | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like Deepgram to enhance your workflow.

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom speech-to-text transcription, audio indexing, or meeting transcription system for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top