AssemblyAI icon
audio-speech

AssemblyAI

Fast batch speech-to-text with rich metadata, strong jargon and mixed-language results, but overlap-heavy meetings can still lose too much.

Visit AssemblyAI
Batch APIWord timestampsSpeaker labels0.016–0.018 RTF
TL;DR — our verdictUpdated September 2026 · 12 test artifacts

Strong batch STT, but crosstalk is the weak spot

Where it wins
  • You need batch transcription that returns completed JSON with word-level timestamps, confidence values, and speaker labels.
  • You transcribe technical narration or jargon-heavy audio and want low WER on those clips.
  • You need mixed-language batch transcription and can accept that this benchmark only covered a mostly-English code-switching sample.
Main limitation
  • You need overlap-heavy meeting audio to retain every word.
Pricing (verified plans)
Free tier $50 in free creditsPay-as-you-go — Universal-3.5 Pro $0.21 / hrPay-as-you-go — Universal-2 $0.15 / hrStreaming — Universal-3.5 Pro Realtime $0.45 / hr
Strongest test artifacts

Feature scores on this page: 33.2/100 (1 scored feature)

Our take

AssemblyAI completed all three batch jobs and stayed consistently fast, returning completed JSON with word-level timestamps, confidence, and speaker labels every time. It was excellent on the medical-jargon clip and best on the bilingual code-switching clip, but the overlapping-speech sample dropped 1,976 words, so it is a solid choice for structured batch transcription as long as overlap-heavy meetings are not the main workload.

Tutorial recording of the AssemblyAI batch transcription benchmark run.

In-Depth Review

Our detailed analysis of AssemblyAI — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Robust Speech-to-Text Transcription
Too much content is lost on crosstalk-heavy audio.
33.16/100
Test Summary
Feature tested: Robust Speech-to-Text Transcription
Result: Failed (33.16/100) — Too much content is lost on crosstalk-heavy audio.

Feature tested: Robust Speech-to-Text Transcription

Result: Failed (33.16/100)

Verdict: Too much content is lost on crosstalk-heavy audio.

Expected behavior: Transcribes difficult speech inputs into text, including overlap-heavy conversations, dense technical narration, and Spanish-English code-switching. The evidence came from the overlap-heavy meeting clip, the Gray's Anatomy clip, and a mostly-English mixed-language sample.

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Input 1: Overlapping Speech / Crosstalk — crosstalk.wav, 65.39 MB, 2142.709 s, AMI EN2002a four-way overlap. — crosstalk.wav

Observed output: Output artifact (Text/code file): The completed transcript for the crosstalk clip finished successfully, but the scored run recorded 33.16% WER with 1,976 deletions and 4 detected speakers, so the transcript dropped too much content for record-quality use. — raw-response.json

Input artifact: Input artifact (Audio file): Input 1: Overlapping Speech / Crosstalk — crosstalk.wav, 65.39 MB, 2142.709 s, AMI EN2002a four-way overlap. — crosstalk.wav

Output artifact: Output artifact (Text/code file): The completed transcript for the crosstalk clip finished successfully, but the scored run recorded 33.16% WER with 1,976 deletions and 4 detected speakers, so the transcript dropped too much content for record-quality use. — raw-response.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Input 2: Medical Jargon — medical_terms.mp3, clean single-speaker narration from Gray's Anatomy with many anatomical terms. — medical_terms.mp3

Observed output: Output artifact (Text/code file): The completed transcript for the medical-jargon clip was highly accurate: the benchmark recorded 3.78% WER and 100% recall of the scored medical terms. — raw-response-2.json

Input artifact: Input artifact (Audio file): Input 2: Medical Jargon — medical_terms.mp3, clean single-speaker narration from Gray's Anatomy with many anatomical terms. — medical_terms.mp3

Output artifact: Output artifact (Text/code file): The completed transcript for the medical-jargon clip was highly accurate: the benchmark recorded 3.78% WER and 100% recall of the scored medical terms. — raw-response-2.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Input 3: Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation with English dominant in the sample. — mix_language.mp3

Observed output: Output artifact (Text/code file): The completed transcript for the bilingual clip was the strongest of the benchmarked engines on this input, with 21.04% WER and 72.5% Spanish token recall. — raw-response-3.json

Input artifact: Input artifact (Audio file): Input 3: Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation with English dominant in the sample. — mix_language.mp3

Output artifact: Output artifact (Text/code file): The completed transcript for the bilingual clip was the strongest of the benchmarked engines on this input, with 21.04% WER and 72.5% Spanish token recall. — raw-response-3.json

What changed: Audio file transformed into Text/code file

Why it matters / Conclusion: It can process overlap-heavy meetings, but the omission rate is high enough that this is not a safe transcript-of-record result.

Transcribes difficult speech inputs into text, including overlap-heavy conversations, dense technical narration, and Spanish-English code-switching. The evidence came from the overlap-heavy meeting clip, the Gray's Anatomy clip, and a mostly-English mixed-language sample.

audio
0:00 / 0:00
Loading audio...
Input 1: Overlapping Speech / Crosstalk — crosstalk.wav, 65.39 MB, 2142.709 s, AMI EN2002a four-way overlap.
OUTPUT
raw-response.json
Loading file...
The completed transcript for the crosstalk clip finished successfully, but the scored run recorded 33.16% WER with 1,976 deletions and 4 detected speakers, so the transcript dropped too much content for record-quality use.
audio
0:00 / 0:00
Loading audio...
Input 2: Medical Jargon — medical_terms.mp3, clean single-speaker narration from Gray's Anatomy with many anatomical terms.
OUTPUT
raw-response-2.json
Loading file...
The completed transcript for the medical-jargon clip was highly accurate: the benchmark recorded 3.78% WER and 100% recall of the scored medical terms.
audio
0:00 / 0:00
Loading audio...
Input 3: Bilingual Code-Switching — mix_language.mp3, spontaneous Spanish-English conversation with English dominant in the sample.
OUTPUT
raw-response-3.json
Loading file...
The completed transcript for the bilingual clip was the strongest of the benchmarked engines on this input, with 21.04% WER and 72.5% Spanish token recall.
Bottom Line
It can process overlap-heavy meetings, but the omission rate is high enough that this is not a safe transcript-of-record result.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Transcript Metadata Export
Consistently available on every run.
Test Summary
Feature tested: Transcript Metadata Export
Result: Passed — Consistently available on every run.

Feature tested: Transcript Metadata Export

Result: Passed

Verdict: Consistently available on every run.

Expected behavior: Returns structured transcript payloads with word-level timestamps, confidence values, and related JSON fields. The evidence shows these metadata fields present consistently across the tested audio runs.

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT: Overlapping Speech / Crosstalk — check returned payload structure and metadata. — crosstalk.wav

Observed output: Output artifact (Text/code file): Raw JSON payload showing word timestamps, confidence values, speaker labels, and 4 distinct speakers detected in the crosstalk run. — raw-response.json

Input artifact: Input artifact (Audio file): INPUT: Overlapping Speech / Crosstalk — check returned payload structure and metadata. — crosstalk.wav

Output artifact: Output artifact (Text/code file): Raw JSON payload showing word timestamps, confidence values, speaker labels, and 4 distinct speakers detected in the crosstalk run. — raw-response.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT: Medical Jargon — check returned payload structure and metadata. — medical_terms.mp3

Observed output: Output artifact (Text/code file): Raw JSON payload showing word timestamps, confidence values, speaker labels, and 1 distinct speaker detected in the medical-jargon run. — raw-response-2.json

Input artifact: Input artifact (Audio file): INPUT: Medical Jargon — check returned payload structure and metadata. — medical_terms.mp3

Output artifact: Output artifact (Text/code file): Raw JSON payload showing word timestamps, confidence values, speaker labels, and 1 distinct speaker detected in the medical-jargon run. — raw-response-2.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT: Bilingual Code-Switching — check returned payload structure and metadata. — mix_language.mp3

Observed output: Output artifact (Text/code file): Raw JSON payload showing word timestamps, confidence values, speaker labels, and 2 distinct speakers detected in the bilingual run. — raw-response-3.json

Input artifact: Input artifact (Audio file): INPUT: Bilingual Code-Switching — check returned payload structure and metadata. — mix_language.mp3

Output artifact: Output artifact (Text/code file): Raw JSON payload showing word timestamps, confidence values, speaker labels, and 2 distinct speakers detected in the bilingual run. — raw-response-3.json

What changed: Audio file transformed into Text/code file

Why it matters / Conclusion: This is a dependable structured-transcript export: timestamps, confidence, and deep JSON payloads were present on every run.

Returns structured transcript payloads with word-level timestamps, confidence values, and related JSON fields. The evidence shows these metadata fields present consistently across the tested audio runs.

audio
0:00 / 0:00
Loading audio...
INPUT: Overlapping Speech / Crosstalk — check returned payload structure and metadata.
OUTPUT
raw-response.json
Loading file...
Raw JSON payload showing word timestamps, confidence values, speaker labels, and 4 distinct speakers detected in the crosstalk run.
audio
0:00 / 0:00
Loading audio...
INPUT: Medical Jargon — check returned payload structure and metadata.
OUTPUT
raw-response-2.json
Loading file...
Raw JSON payload showing word timestamps, confidence values, speaker labels, and 1 distinct speaker detected in the medical-jargon run.
audio
0:00 / 0:00
Loading audio...
INPUT: Bilingual Code-Switching — check returned payload structure and metadata.
OUTPUT
raw-response-3.json
Loading file...
Raw JSON payload showing word timestamps, confidence values, speaker labels, and 2 distinct speakers detected in the bilingual run.
Bottom Line
This is a dependable structured-transcript export: timestamps, confidence, and deep JSON payloads were present on every run.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Speaker Diarization
Labels are detected, but attribution correctness was not measured.
Test Summary
Feature tested: Speaker Diarization
Result: Partial — Labels are detected, but attribution correctness was not measured.

Feature tested: Speaker Diarization

Result: Partial

Verdict: Labels are detected, but attribution correctness was not measured.

Expected behavior: Adds speaker labels or speaker IDs to transcript output and reports the number of distinct speakers detected. It was exercised on crosstalk, medical-jargon, and bilingual clips, where speaker labels were surfaced consistently.

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Input 1: Overlapping Speech / Crosstalk — four participants speaking over one another. — crosstalk.wav

Observed output: Output artifact (Text/code file): The transcript payload reported 4 distinct speaker labels on the crosstalk clip, matching the four-participant meeting structure, but label correctness was not scored. — raw-response.json

Input artifact: Input artifact (Audio file): Input 1: Overlapping Speech / Crosstalk — four participants speaking over one another. — crosstalk.wav

Output artifact: Output artifact (Text/code file): The transcript payload reported 4 distinct speaker labels on the crosstalk clip, matching the four-participant meeting structure, but label correctness was not scored. — raw-response.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Input 2: Medical Jargon — single-speaker narration. — medical_terms.mp3

Observed output: Output artifact (Text/code file): The transcript payload reported 1 detected speaker on the medical-jargon clip, consistent with the single-speaker narration. — raw-response-2.json

Input artifact: Input artifact (Audio file): Input 2: Medical Jargon — single-speaker narration. — medical_terms.mp3

Output artifact: Output artifact (Text/code file): The transcript payload reported 1 detected speaker on the medical-jargon clip, consistent with the single-speaker narration. — raw-response-2.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Input 3: Bilingual Code-Switching — two-speaker Spanish-English conversation. — mix_language.mp3

Observed output: Output artifact (Text/code file): The transcript payload reported 2 detected speakers on the bilingual clip, but speaker-attribution correctness was not independently measured. — raw-response-3.json

Input artifact: Input artifact (Audio file): Input 3: Bilingual Code-Switching — two-speaker Spanish-English conversation. — mix_language.mp3

Output artifact: Output artifact (Text/code file): The transcript payload reported 2 detected speakers on the bilingual clip, but speaker-attribution correctness was not independently measured. — raw-response-3.json

What changed: Audio file transformed into Text/code file

Why it matters / Conclusion: Speaker labels are exposed consistently, but this benchmark did not verify whether every attribution was correct.

Adds speaker labels or speaker IDs to transcript output and reports the number of distinct speakers detected. It was exercised on crosstalk, medical-jargon, and bilingual clips, where speaker labels were surfaced consistently.

audio
0:00 / 0:00
Loading audio...
Input 1: Overlapping Speech / Crosstalk — four participants speaking over one another.
OUTPUT
raw-response.json
Loading file...
The transcript payload reported 4 distinct speaker labels on the crosstalk clip, matching the four-participant meeting structure, but label correctness was not scored.
audio
0:00 / 0:00
Loading audio...
Input 2: Medical Jargon — single-speaker narration.
OUTPUT
raw-response-2.json
Loading file...
The transcript payload reported 1 detected speaker on the medical-jargon clip, consistent with the single-speaker narration.
audio
0:00 / 0:00
Loading audio...
Input 3: Bilingual Code-Switching — two-speaker Spanish-English conversation.
OUTPUT
raw-response-3.json
Loading file...
The transcript payload reported 2 detected speakers on the bilingual clip, but speaker-attribution correctness was not independently measured.
Bottom Line
Speaker labels are exposed consistently, but this benchmark did not verify whether every attribution was correct.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Batch Transcription
Completed successfully on all three test runs.
Test Summary
Feature tested: Batch Transcription
Result: Partial — Completed successfully on all three test runs.

Feature tested: Batch Transcription

Result: Partial

Verdict: Completed successfully on all three test runs.

Expected behavior: Accepts uploaded audio files, runs them through a batch transcription flow, and returns a completed transcript JSON response. The evidence covers overlapping speech, medical narration, and bilingual code-switching inputs reaching completion through the API path.

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT: Overlapping Speech / Crosstalk — crosstalk.wav (65.39 MB, 2142.709s) — crosstalk.wav

Observed output: Output artifact (Text/code file): Completed JSON response for the crosstalk run; the job reached status completed and returned a transcript payload. — raw-response.json

Input artifact: Input artifact (Audio file): INPUT: Overlapping Speech / Crosstalk — crosstalk.wav (65.39 MB, 2142.709s) — crosstalk.wav

Output artifact: Output artifact (Text/code file): Completed JSON response for the crosstalk run; the job reached status completed and returned a transcript payload. — raw-response.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT: Medical Jargon — medical_terms.mp3 (8.58 MB, 1123.944s) — medical_terms.mp3

Observed output: Output artifact (Text/code file): Completed JSON response for the medical-jargon run; the job reached status completed and returned a transcript payload. — raw-response-2.json

Input artifact: Input artifact (Audio file): INPUT: Medical Jargon — medical_terms.mp3 (8.58 MB, 1123.944s) — medical_terms.mp3

Output artifact: Output artifact (Text/code file): Completed JSON response for the medical-jargon run; the job reached status completed and returned a transcript payload. — raw-response-2.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT: Bilingual Code-Switching — mix_language.mp3 (22.19 MB, 1938.495s) — mix_language.mp3

Observed output: Output artifact (Text/code file): Completed JSON response for the bilingual run; the job reached status completed and returned a transcript payload. — raw-response-3.json

Input artifact: Input artifact (Audio file): INPUT: Bilingual Code-Switching — mix_language.mp3 (22.19 MB, 1938.495s) — mix_language.mp3

Output artifact: Output artifact (Text/code file): Completed JSON response for the bilingual run; the job reached status completed and returned a transcript payload. — raw-response-3.json

What changed: Audio file transformed into Text/code file

Why it matters / Conclusion: The batch API path was reliable across all three runs, with no manual intervention needed to reach completed JSON.

Accepts uploaded audio files, runs them through a batch transcription flow, and returns a completed transcript JSON response. The evidence covers overlapping speech, medical narration, and bilingual code-switching inputs reaching completion through the API path.

audio
0:00 / 0:00
Loading audio...
INPUT: Overlapping Speech / Crosstalk — crosstalk.wav (65.39 MB, 2142.709s)
OUTPUT
raw-response.json
Loading file...
Completed JSON response for the crosstalk run; the job reached status completed and returned a transcript payload.
audio
0:00 / 0:00
Loading audio...
INPUT: Medical Jargon — medical_terms.mp3 (8.58 MB, 1123.944s)
OUTPUT
raw-response-2.json
Loading file...
Completed JSON response for the medical-jargon run; the job reached status completed and returned a transcript payload.
audio
0:00 / 0:00
Loading audio...
INPUT: Bilingual Code-Switching — mix_language.mp3 (22.19 MB, 1938.495s)
OUTPUT
raw-response-3.json
Loading file...
Completed JSON response for the bilingual run; the job reached status completed and returned a transcript payload.
Bottom Line
The batch API path was reliable across all three runs, with no manual intervention needed to reach completed JSON.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research

How it scored on the research's own criteria

The 3 evaluation dimensions from our hands-on research on AssemblyAI, each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Output qualityMixed3/5Accuracy is strong on clean medical narration, but it drops a lot on overlapping speech and settles in the middle on mixed-language audio. Because the best case is excellent while the hardest case is weak, the overall quality picture is mixed rather than consistently strong.open proof ↗
Automation levelStrong5/5Each run completed the full upload → transcription → polling flow on its own and returned a finished result without operator intervention. The only missing detail is per-call timing, not whether the workflow itself ran end to end, so this stays at the top score.open proof ↗
Input handlingStrong5/5It accepted all three uploaded files without objection and finished each run quickly, with low real-time factors around 0.016–0.018 and modest costs. That combination of reliable acceptance plus consistently fast batch processing supports the top score.open proof ↗

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Official pricing

Benchmark runs used Universal-3.5 Pro; diarization stacks on top as an add-on.

Free tier
$50 in free credits
No credit card required; 5 new streaming connections/min.
TESTED
Pay-as-you-go — Universal-3.5 Pro
$0.21 / hr
100 new streams/min; most accurate async model; 18 languages, native code switching.
Pay-as-you-go — Universal-2
$0.15 / hr
99 languages; "exceptional accuracy at a lower price."
Streaming — Universal-3.5 Pro Realtime
$0.45 / hr
Billed on WebSocket session duration, not audio duration.

The report says add-ons stack on top of the base rate, and multichannel audio is billed per channel.

✓ Use This If
You need batch transcription that returns completed JSON with word-level timestamps, confidence values, and speaker labels.
You transcribe technical narration or jargon-heavy audio and want low WER on those clips.
You need mixed-language batch transcription and can accept that this benchmark only covered a mostly-English code-switching sample.
✕ Skip This If
You need overlap-heavy meeting audio to retain every word.
You need validated speaker-attribution correctness rather than detected speaker labels.
You need streaming or live latency from this evaluation; it was not measured.
You need a balanced bilingual routing test; input-3 was mostly English.
audio-speechaudio-to-texttextOther
On the overlapping-speech / crosstalk clip, it scored 33.16% WER with 461 substitutions, 1,976 deletions, and 76 insertions against a 7,579-word reference. The run completed and detected 4 speaker labels, but too much content was omitted for transcript-of-record use.
It performed very well on the medical-jargon clip: 3.78% WER, 71 substitutions, 14 deletions, and 18 insertions against a 2,728-word reference. The transcript detail also reported 100% recall of the scored jargon terms.
Yes. On the bilingual code-switching clip it scored 21.04% WER and 72.5% Spanish token recall, and it returned a completed transcript without special configuration. The report also notes that the sample was mostly English, so balanced bilingual routing was not directly stress-tested.
The raw responses showed word-level timestamps, confidence values, and speaker labels. The payload depth was reported as 3/3, with deep JSON nesting and distinct speaker keys present in every run.
No. This was a batch benchmark, and streaming latency was explicitly not measured.
The report lists a $50 free tier in credits, Universal-3.5 Pro at $0.21/hr, Universal-2 at $0.15/hr, and Universal-3.5 Pro Realtime at $0.45/hr. It also says add-ons stack separately, including async diarization at +$0.02/hr.

Banner Preview

How the embed badge will look on your site

AssemblyAI featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/assemblyai-speech-to-text?utm_source=assemblyai-speech-to-text_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="AssemblyAI | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like AssemblyAI to enhance your workflow.

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom speech-to-text, transcription, or meeting transcription system for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top