Deepgram
Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching.
Mixed results, strong metadata
- you need a batch STT API that returns word-level timing, confidence, and speaker labels
- you primarily transcribe single-language technical narration and care about jargon recall
- you can process audio offline and care more about throughput than live streaming
- you need strong Spanish-English code-switching or multilingual transcription
Our take
Deepgram Nova-3 returned a full developer payload on every run and was excellent on the medical-jargon clip, but the crosstalk and bilingual runs were noisy enough that I would not treat it as reliable for overlapping or code-switched audio as configured here.
In-Depth Review
Our detailed analysis of Deepgram — features, performance, and real-world testing.
Feature-by-Feature Breakdown
Audio TranscriptionMixed▾
Feature tested: Audio Transcription
Result: Failed
Verdict: Mixed
Expected behavior: Deepgram converts uploaded pre-recorded spoken audio into text through a single API request. The member cards exercised it on benchmark clips including overlapping meetings, medical-jargon narration, Spanish-English code-switching, and other prerecorded files.
Test case: Image → Text/code file
Input type: Image
Input used: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, 65.39 MB, 2142.709 s, four-way overlapping meeting audio. — 84e7dcb47c124944ad4547c508cf5ee5.png
Observed output: Output artifact (Text/code file): Crosstalk run: Deepgram returned a structured JSON transcript for crosstalk.wav; the scored run posted 36.27% WER, 63.87s latency, 0.02981 RTF, 6,946 returned words, and 4 detected speakers. — raw-response.json
Input artifact: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, 65.39 MB, 2142.709 s, four-way overlapping meeting audio. — 84e7dcb47c124944ad4547c508cf5ee5.png
Output artifact: Output artifact (Text/code file): Crosstalk run: Deepgram returned a structured JSON transcript for crosstalk.wav; the scored run posted 36.27% WER, 63.87s latency, 0.02981 RTF, 6,946 returned words, and 4 detected speakers. — raw-response.json
What changed: Image transformed into Text/code file
Test case: Image → Text/code file
Input type: Image
Input used: Input artifact (Image): Medical Jargon — medical_terms.mp3, 18:44, 8.58 MB, 1123.971 s, single-speaker narration dense with anatomical terms. — 3f46ad25af044fcea47435b057d24056.png
Observed output: Output artifact (Text/code file): Medical-jargon run: Deepgram returned a transcript JSON for medical_terms.mp3; the scored run posted 5.43% WER, 12.02s latency, 0.01069 RTF, 2,726 returned words against 2,728 reference words, and 100% jargon recall. — raw-response-2.json
Input artifact: Input artifact (Image): Medical Jargon — medical_terms.mp3, 18:44, 8.58 MB, 1123.971 s, single-speaker narration dense with anatomical terms. — 3f46ad25af044fcea47435b057d24056.png
Output artifact: Output artifact (Text/code file): Medical-jargon run: Deepgram returned a transcript JSON for medical_terms.mp3; the scored run posted 5.43% WER, 12.02s latency, 0.01069 RTF, 2,726 returned words against 2,728 reference words, and 100% jargon recall. — raw-response-2.json
What changed: Image transformed into Text/code file
Test case: Image → Text/code file
Input type: Image
Input used: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, 32:18, 22.19 MB, 1938.495 s, spontaneous Spanish-English conversation. — 353dc535361142fbab8920650fb7d1ca.png
Observed output: Output artifact (Text/code file): Bilingual run: Deepgram returned a transcript JSON for mix_language.mp3; the scored run posted 38.13% WER, 27.58s latency, 0.01423 RTF, 5,691 returned words against 6,517 reference words, and Spanish recall was only 3.8%. — raw-response-3.json
Input artifact: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, 32:18, 22.19 MB, 1938.495 s, spontaneous Spanish-English conversation. — 353dc535361142fbab8920650fb7d1ca.png
Output artifact: Output artifact (Text/code file): Bilingual run: Deepgram returned a transcript JSON for mix_language.mp3; the scored run posted 38.13% WER, 27.58s latency, 0.01423 RTF, 5,691 returned words against 6,517 reference words, and Spanish recall was only 3.8%. — raw-response-3.json
What changed: Image transformed into Text/code file
Test case: Image → Image
Input type: Image
Input used: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, batch throughput and latency run. — 84e7dcb47c124944ad4547c508cf5ee5.png
Observed output: Output artifact (Image): Run metrics for crosstalk.wav show 36.27% WER, 63.87s wall-clock latency, 0.02981 RTF, and an estimated cost of $0.22498 for the batch run. — 03-terminal-metrics-input-1.png
Input artifact: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, batch throughput and latency run. — 84e7dcb47c124944ad4547c508cf5ee5.png
Output artifact: Output artifact (Image): Run metrics for crosstalk.wav show 36.27% WER, 63.87s wall-clock latency, 0.02981 RTF, and an estimated cost of $0.22498 for the batch run. — 03-terminal-metrics-input-1.png
What changed: Image transformed into Image
Test case: Image → Image
Input type: Image
Input used: Input artifact (Image): Medical Jargon — medical_terms.mp3, batch throughput and latency run. — 3f46ad25af044fcea47435b057d24056.png
Observed output: Output artifact (Image): Run metrics for medical_terms.mp3 show 5.43% WER, 12.02s wall-clock latency, 0.01069 RTF, and an estimated cost of $0.11802 for the batch run. — 03-terminal-metrics-input-2.png
Input artifact: Input artifact (Image): Medical Jargon — medical_terms.mp3, batch throughput and latency run. — 3f46ad25af044fcea47435b057d24056.png
Output artifact: Output artifact (Image): Run metrics for medical_terms.mp3 show 5.43% WER, 12.02s wall-clock latency, 0.01069 RTF, and an estimated cost of $0.11802 for the batch run. — 03-terminal-metrics-input-2.png
What changed: Image transformed into Image
Test case: Image → Image
Input type: Image
Input used: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, batch throughput and latency run. — 353dc535361142fbab8920650fb7d1ca.png
Observed output: Output artifact (Image): Run metrics for mix_language.mp3 show 38.13% WER, 27.58s wall-clock latency, 0.01423 RTF, and an estimated cost of $0.20354 for the batch run. — 03-terminal-metrics-input-3.png
Input artifact: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, batch throughput and latency run. — 353dc535361142fbab8920650fb7d1ca.png
Output artifact: Output artifact (Image): Run metrics for mix_language.mp3 show 38.13% WER, 27.58s wall-clock latency, 0.01423 RTF, and an estimated cost of $0.20354 for the batch run. — 03-terminal-metrics-input-3.png
What changed: Image transformed into Image
Why it matters / Conclusion: Useful for batch transcription when the audio is mostly single-speaker and monolingual, but not reliable for crosstalk-heavy meetings or code-switched conversation.
Deepgram converts uploaded pre-recorded spoken audio into text through a single API request. The member cards exercised it on benchmark clips including overlapping meetings, medical-jargon narration, Spanish-English code-switching, and other prerecorded files.









Structured Transcript OutputConsistent▾
Feature tested: Structured Transcript Output
Result: Passed
Verdict: Consistent
Expected behavior: Deepgram returns transcript payloads as structured JSON with fields such as word-level timing, confidence values, speaker labels, punctuated tokens, and other parseable metadata. The member cards exercised this output shape across runs and options like smart_format, diarize, punctuate, and utterances.
Test case: Image → Image
Input type: Image
Input used: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, batch transcription run used to inspect JSON structure and metadata. — 84e7dcb47c124944ad4547c508cf5ee5.png
Observed output: Output artifact (Image): Raw API response excerpt for crosstalk.wav showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers. — 02-response-raw-input-1.png
Input artifact: Input artifact (Image): Overlapping Speech / Crosstalk — crosstalk.wav, batch transcription run used to inspect JSON structure and metadata. — 84e7dcb47c124944ad4547c508cf5ee5.png
Output artifact: Output artifact (Image): Raw API response excerpt for crosstalk.wav showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers. — 02-response-raw-input-1.png
What changed: Image transformed into Image
Test case: Image → Image
Input type: Image
Input used: Input artifact (Image): Medical Jargon — medical_terms.mp3, batch transcription run used to inspect JSON structure and metadata. — 3f46ad25af044fcea47435b057d24056.png
Observed output: Output artifact (Image): Raw API response excerpt for medical_terms.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 5,845 timed tokens, and 1 distinct speaker. — 02-response-raw-input-2.png
Input artifact: Input artifact (Image): Medical Jargon — medical_terms.mp3, batch transcription run used to inspect JSON structure and metadata. — 3f46ad25af044fcea47435b057d24056.png
Output artifact: Output artifact (Image): Raw API response excerpt for medical_terms.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 5,845 timed tokens, and 1 distinct speaker. — 02-response-raw-input-2.png
What changed: Image transformed into Image
Test case: Image → Image
Input type: Image
Input used: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, batch transcription run used to inspect JSON structure and metadata. — 353dc535361142fbab8920650fb7d1ca.png
Observed output: Output artifact (Image): Raw API response excerpt for mix_language.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 11,684 timed tokens, and 3 distinct speakers. — 02-response-raw-input-3.png
Input artifact: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, batch transcription run used to inspect JSON structure and metadata. — 353dc535361142fbab8920650fb7d1ca.png
Output artifact: Output artifact (Image): Raw API response excerpt for mix_language.mp3 showing metadata, model info, the opening transcript, and developer-feature detection: word_timestamps yes, confidence yes, speaker_labels yes, 11,684 timed tokens, and 3 distinct speakers. — 02-response-raw-input-3.png
What changed: Image transformed into Image
Why it matters / Conclusion: This is the most dependable part of the product in this benchmark: the output shape stayed consistent enough for downstream parsing on every clip.
Deepgram returns transcript payloads as structured JSON with fields such as word-level timing, confidence values, speaker labels, punctuated tokens, and other parseable metadata. The member cards exercised this output shape across runs and options like smart_format, diarize, punctuate, and utterances.






Speaker DiarizationPartial▾
Feature tested: Speaker Diarization
Result: Partial
Verdict: Partial
Expected behavior: Deepgram can surface speaker labels in its transcript output and, on the hardest crosstalk clip, matched the four-participant count. The benchmark cards note that speaker attribution quality itself was not fully verified.
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a). — crosstalk.wav
Observed output: Output artifact (Image): The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference. — 03-terminal-metrics.png
Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk test clip: crosstalk.wav (35:43, 65.39 MB, AMI EN2002a). — crosstalk.wav
Output artifact: Output artifact (Image): The run metrics reported distinct_speaker_labels 4 on the overlap clip, which matched the four participants in the benchmark reference. — 03-terminal-metrics.png
What changed: Audio file transformed into Image
Test case: Audio file → Image
Input type: Audio file
Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk — crosstalk.wav, four-way overlapping meeting audio with diarization stress. — crosstalk.wav
Observed output: Output artifact (Image): Transcript detail for the crosstalk clip shows the largest divergence, but also confirms diarization detection with four labels and four true participants; the transcript itself was still noisy, with a 36.27% WER and 510 insertions. — 04-transcript-detail-input-1.png
Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk — crosstalk.wav, four-way overlapping meeting audio with diarization stress. — crosstalk.wav
Output artifact: Output artifact (Image): Transcript detail for the crosstalk clip shows the largest divergence, but also confirms diarization detection with four labels and four true participants; the transcript itself was still noisy, with a 36.27% WER and 510 insertions. — 04-transcript-detail-input-1.png
What changed: Audio file transformed into Image
Test case: Image → Image
Input type: Image
Input used: Input artifact (Image): Medical Jargon — medical_terms.mp3, single-speaker narration used to verify speaker-label detection. — 3f46ad25af044fcea47435b057d24056.png
Observed output: Output artifact (Image): Raw response excerpt for the medical-jargon clip shows speaker_labels yes and a single distinct speaker, which is consistent with the single-speaker source audio. — 02-response-raw-input-2.png
Input artifact: Input artifact (Image): Medical Jargon — medical_terms.mp3, single-speaker narration used to verify speaker-label detection. — 3f46ad25af044fcea47435b057d24056.png
Output artifact: Output artifact (Image): Raw response excerpt for the medical-jargon clip shows speaker_labels yes and a single distinct speaker, which is consistent with the single-speaker source audio. — 02-response-raw-input-2.png
What changed: Image transformed into Image
Test case: Image → Image
Input type: Image
Input used: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation used to verify speaker-label detection. — 353dc535361142fbab8920650fb7d1ca.png
Observed output: Output artifact (Image): Raw response excerpt for the bilingual clip shows speaker_labels yes and three distinct speakers, but the benchmark did not measure whether those labels were attributed to the right voices. — 02-response-raw-input-3.png
Input artifact: Input artifact (Image): Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation used to verify speaker-label detection. — 353dc535361142fbab8920650fb7d1ca.png
Output artifact: Output artifact (Image): Raw response excerpt for the bilingual clip shows speaker_labels yes and three distinct speakers, but the benchmark did not measure whether those labels were attributed to the right voices. — 02-response-raw-input-3.png
What changed: Image transformed into Image
Why it matters / Conclusion: Good for speaker-count metadata, but attribution quality remains unproven.
Deepgram can surface speaker labels in its transcript output and, on the hardest crosstalk clip, matched the four-participant count. The benchmark cards note that speaker attribution quality itself was not fully verified.






How it scored on the research's own criteria
The 3 evaluation dimensions from our hands-on research on Deepgram , each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.
held up partial failed not exercised by this input
| Criterion | Verdict | What the runs showed | Per input | Proof |
|---|---|---|---|---|
| Output quality | Weak2/5 | The model is excellent on clean medical narration, but the other two runs are badly hurt by heavy error rates, and the bilingual case especially shows that it loses most Spanish even when the conversation is mostly English. | open proof ↗ | |
| Automation level | Strong5/5 | Each run completed as a single automated request-response flow with no manual intervention, so the workflow is as hands-off as it can be in this setup. | — | |
| Input handling | Strong5/5 | It handled all three audio files cleanly, stayed fast on each run, and kept costs modest, so this is consistently top-tier input handling rather than just passing the basics. | open proof ↗ |
Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.
Official pricing from the vendor page
The benchmark config was logged at $0.0063/min ($0.378/audio-hour); the report notes that the served pre-recorded Nova-3 tab could not be fully verified in-browser.
The report also records Nova-3 Monolingual at $0.0048/min and Nova-3 Multilingual at $0.0058/min, with Speaker Diarization listed separately at $0.0020/min. Smart Formatting is included.
Banner Preview
How the embed badge will look on your site

Embed HTML
Copy this code to your website source
Quick Integration Guide
- 1Copy the HTML code block above.
- 2Paste it into your site's HTML or CMS editor.
- 3Banner appears instantly on your page.
- 4Links back to your tool profile here.
Similar Tools
Discover more AI tools like Deepgram to enhance your workflow.
Comments (0)
Need a custom AI solution for this use case?
If you are looking to build a custom speech-to-text transcription, audio indexing, or meeting transcription system for your business or internal workflow, email us at contact@futuresmart.ai.
Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.