--- title: "ElevenLabs" type: "AI Tool" url: "https://aidemos.com/tools/elevenlabs-scribe" description: "ElevenLabs Scribe returned word timestamps, confidence, and speaker labels on jargon audio, but crosstalk and Spanish code-switching caused errors." category: "audio-speech" website: "https://elevenlabs.io/docs/api-reference/speech-to-text" published: "2026-08-13T09:18:22.530033+00:00" updated: "2026-09-01T02:54:52.298943+00:00" evidenceCount: 33 verifiedCount: 26 coverage: "dense" --- # ElevenLabs Fast batch speech-to-text with word-level metadata, strongest on jargon and weaker on overlap/code-switching. ## TL;DR Verdict **Strong default batch STT API** **Where it wins:** - You need batch speech-to-text from a single multipart POST. - You need word-level timestamps, confidence values, and speaker labels in the transcript payload. - You need strong accuracy on jargon-heavy English narration. **Main limitation:** You need validated speaker attribution correctness on overlapping conversations. **Pricing:** Free / Pay-as-you-go $0.22/hr · Starter $6/month · Creator $22/month (first month $11) · Pro $99/month `Word timestamps` · `Speaker labels` · `3 scored runs` · `Fast batch` **Website:** [Visit ElevenLabs](https://elevenlabs.io/docs/api-reference/speech-to-text) ## Evidence (first-party, tested) *33 tested cells · 26/33 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:elevenlabs·cross·automation-level`.* | Criterion | Scenario | Verdict | Score | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | | Automation level | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-317850cd4009.md) | `ev:elevenlabs·cross·automation-level` | | Control Granularity | High-Quality Voice Sample | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/016c6dfc3e444b4ebb9c178a5a7a74c9.wav?v=1) | `ev:elevenlabs·high-quality-voice-sample·control-granularity` | | Control Granularity | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7e568a7c0b154543ad8280ce25f788ea.png?v=1) | `ev:elevenlabs·cross·control-granularity` | | Control Granularity | Multilingual Voice Sample (Hindi) | ◐ mixed | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-input-multilingual-77a768fa7ef3.txt) | `ev:elevenlabs·multilingual-voice-sample-hindi·control-granularity` | | Control Granularity | Low-Quality Voice Sample | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/1d29cd4de48647b68942f0a3137e700a.wav?v=1) | `ev:elevenlabs·low-quality-voice-sample·control-granularity` | | Export | cross-scenario | ✓ worked | — | 👁 observed | `ev:elevenlabs·cross·export` | | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/94df22cc6deb43509f601924e291331f.png?v=1) | `ev:elevenlabs·bilingual-spanish-english-code-switching-speech·export` | | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/afa813794c9f4223ae0776d92c91ad7b.png?v=1) | `ev:elevenlabs·medical-anatomy-narration-with-dense-jargon·export` | | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4e0fd19a33914c36a0a75146eb9accc4.png?v=1) | `ev:elevenlabs·overlapping-meeting-speech-with-cross-talk·export` | | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5c37964b48fe4b0faa889b7a07f04b08.png?v=1) | `ev:elevenlabs·medical-anatomy-narration-with-dense-jargon·input-handling` | | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-317850cd4009.md) | `ev:elevenlabs·cross·input-handling` | | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/545e483d71254ddba884f07d49c5e714.png?v=1) | `ev:elevenlabs·overlapping-meeting-speech-with-cross-talk·input-handling` | | Input handling | Bilingual Spanish-English code-switching speech | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/48d19a5d4d364648bf535b1295d8290f.png?v=1) | `ev:elevenlabs·bilingual-spanish-english-code-switching-speech·input-handling` | | Long-Form Consistency | Low-Quality Voice Sample | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/1d29cd4de48647b68942f0a3137e700a.wav?v=1) | `ev:elevenlabs·low-quality-voice-sample·long-form-consistency` | | Long-Form Consistency | High-Quality Voice Sample | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/016c6dfc3e444b4ebb9c178a5a7a74c9.wav?v=1) | `ev:elevenlabs·high-quality-voice-sample·long-form-consistency` | | Long-Form Consistency | cross-scenario | ◐ mixed | — | 👁 observed | `ev:elevenlabs·cross·long-form-consistency` | | Long-Form Consistency | Multilingual Voice Sample (Hindi) | ⚠ struggled | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-hindi-script-input-8cffd5de94b9.txt) | `ev:elevenlabs·multilingual-voice-sample-hindi·long-form-consistency` | | Multilingual Output Quality | Multilingual Voice Sample (Hindi) | ⚠ struggled | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-hindi-script-input-8cffd5de94b9.txt) | `ev:elevenlabs·multilingual-voice-sample-hindi·multilingual-output-quality` | | Multilingual Output Quality | Low-Quality Voice Sample | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/1d29cd4de48647b68942f0a3137e700a.wav?v=1) | `ev:elevenlabs·low-quality-voice-sample·multilingual-output-quality` | | Multilingual Output Quality | High-Quality Voice Sample | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/016c6dfc3e444b4ebb9c178a5a7a74c9.wav?v=1) | `ev:elevenlabs·high-quality-voice-sample·multilingual-output-quality` | | Multilingual Output Quality | cross-scenario | ◐ mixed | — | 👁 observed | `ev:elevenlabs·cross·multilingual-output-quality` | | Naturalness & Human Quality | Low-Quality Voice Sample | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/1d29cd4de48647b68942f0a3137e700a.wav?v=1) | `ev:elevenlabs·low-quality-voice-sample·naturalness-and-human-quality` | | Naturalness & Human Quality | cross-scenario | ◐ mixed | — | 👁 observed | `ev:elevenlabs·cross·naturalness-and-human-quality` | | Naturalness & Human Quality | High-Quality Voice Sample | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/016c6dfc3e444b4ebb9c178a5a7a74c9.wav?v=1) | `ev:elevenlabs·high-quality-voice-sample·naturalness-and-human-quality` | | Naturalness & Human Quality | Multilingual Voice Sample (Hindi) | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-hindi-script-input-8cffd5de94b9.txt) | `ev:elevenlabs·multilingual-voice-sample-hindi·naturalness-and-human-quality` | | Output quality | Overlapping meeting speech with cross-talk | ⚠ struggled | 26.7/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/b452a213d5ae41e7aaaa6b91c64d4d2d.png?v=1) | `ev:elevenlabs·overlapping-meeting-speech-with-cross-talk·output-quality` | | Output quality | Bilingual Spanish-English code-switching speech | ✗ failed | 28.6/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/51c4450c1717464a9bb49a9a69f0b94d.png?v=1) | `ev:elevenlabs·bilingual-spanish-english-code-switching-speech·output-quality` | | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 3.0/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4128c2a5389e48fea7c1741621f277c6.png?v=1) | `ev:elevenlabs·medical-anatomy-narration-with-dense-jargon·output-quality` | | Output quality | cross-scenario | ◐ mixed | — | 👁 observed | `ev:elevenlabs·cross·output-quality` | | Voice Match Accuracy | High-Quality Voice Sample | ⚠ struggled | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/016c6dfc3e444b4ebb9c178a5a7a74c9.wav?v=1) | `ev:elevenlabs·high-quality-voice-sample·voice-match-accuracy` | | Voice Match Accuracy | Low-Quality Voice Sample | ⚠ struggled | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/1d29cd4de48647b68942f0a3137e700a.wav?v=1) | `ev:elevenlabs·low-quality-voice-sample·voice-match-accuracy` | | Voice Match Accuracy | cross-scenario | ⚠ struggled | — | 👁 observed | `ev:elevenlabs·cross·voice-match-accuracy` | | Voice Match Accuracy | Multilingual Voice Sample (Hindi) | ✗ failed | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-hindi-script-input-8cffd5de94b9.txt) | `ev:elevenlabs·multilingual-voice-sample-hindi·voice-match-accuracy` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. > **Strong default batch STT API** > > Tested here as \`scribe_v1\`, ElevenLabs Scribe looks like a strong batch speech-to-text API for developers: every scored run returned word-level timestamps, confidence values, and speaker labels, and the medical-jargon sample came back at 3.01% WER. The tradeoff is clear on harder audio: crosstalk produced heavy insertions and over-segmentation, and the bilingual run lost Spanish tokens, so it is best when you need a fast, metadata-rich transcript and can tolerate weaker overlap and code-switching. ## Demo Recording [Video: ElevenLabs demo recording](https://cdn.futuresmart.ai/public/aidemos/86a430d970e943d4bec81d7aa92e8984.mov?v=1) *Video — Tutorial recording of the ElevenLabs Scribe benchmark workflow and result review.* ## Feature-by-Feature Breakdown ### Batch Speech-to-Text Transcription **Verdict:** Strong on jargon, weak on overlap/code-switching ElevenLabs Scribe performs one-shot batch transcription over long audio via a single multipart POST. The benchmark exercised it on jargon-heavy narration, crosstalk, and bilingual code-switching, showing the core transcript engine works across these audio inputs but with varying quality. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** The core transcript engine is very input-sensitive: it can be excellent on technical English, but hard conversational audio and code-switching both degrade sharply. ### Transcript Metadata Output **Verdict:** Consistent developer payload The API returns structured transcript payloads with developer-facing metadata such as word-level timestamps, confidence/logprob values, speaker IDs, and stable JSON shape. The evidence came from transcript outputs used for captioning, search, and speaker-aware downstream workflows. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** The payload shape is stable across runs and includes the metadata a downstream developer would expect from a production STT API. ### Speaker Diarization **Verdict:** Present, but attribution correctness unvalidated The engine emits speaker labels and speaker counts in the transcript output, enabling rough diarization and inspection of how many speakers appear in the audio. The benchmark only verified label presence and count, not correct attribution. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Speaker labels are there, but the benchmark only proves label presence and count, not correct attribution. ### Multilingual Transcription ElevenLabs Scribe can transcribe mixed Spanish-English or code-switched audio without special configuration. The sampled evidence shows usable mixed-language transcription, though Spanish recall is weaker than in clean-English narration. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** It can transcribe mixed-language audio, but Spanish quality is materially weaker than the clean-English medical sample. ## Official pricing Current Scribe v2 list pricing from ElevenLabs' pricing page; the benchmark run itself used `scribe_v1`. | Plan | Price | Notes | | --- | --- | --- | | Free / Pay-as-you-go | $0.22/hr | 4 h 30 min of Scribe v2 included; Scribe v2 Realtime is $0.39/hr with 2 h 30 min included. | | Starter | $6/month | $0.22/hr with 27 h of Scribe v2 included; 15 h of Realtime included. | | Creator | $22/month (first month $11) | $0.22/hr with 100 h of Scribe v2 included; 56 h of Realtime included. | | Pro | $99/month | $0.22/hr with 450 h of Scribe v2 included; 254 h of Realtime included. | | Scale | $299/month | $0.22/hr with 1,359 h of Scribe v2 included; 767 h of Realtime included. | | Business | $990/month | $0.22/hr with 4,500 h of Scribe v2 included; 2,538 h of Realtime included. | | Enterprise | Custom | Custom DPA/SLA, SSO, and HIPAA BAA; requires sales contact. | | Startup Grants Program | Free for 12 months | 33,000,000 characters; application required. | *The per-hour Scribe rate is flat across plans at $0.22/hr; paid tiers mainly increase included hours and concurrency. Add-ons are priced separately: entity detection +$0.070/hr and keyterm prompting +$0.050/hr. Prices exclude taxes.* ## Is It Right For You? **Use it if** - You need batch speech-to-text from a single multipart POST. - You need word-level timestamps, confidence values, and speaker labels in the transcript payload. - You need strong accuracy on jargon-heavy English narration. - You need fast batch turnaround at about $0.2202 per audio-hour. **Skip it if** - You need validated speaker attribution correctness on overlapping conversations. - You need overlap-heavy audio to stay faithful without insertions and over-segmentation. - You need bilingual Spanish-English audio to stay near the clean-English result. - You need streaming-latency evidence from this benchmark. ## Classification - **Category:** audio-speech - **Subcategory:** audio-to-text - **Type:** speech - **Built for:** Other ## Frequently Asked Questions **Q: How accurate was ElevenLabs Scribe on overlapping speech?** On the crosstalk sample it scored 26.67% WER, with 856 substitutions, 782 deletions, and 383 insertions. The transcript-detail view shows a large missing span, and the metrics report says it detected 5 speaker labels for a 4-participant meeting. **Q: How did it perform on medical jargon?** It was strongest on the medical-jargon sample: 3.01% WER, 51 substitutions, 8 deletions, and 23 insertions. The benchmark also reports 100.0% jargon recall on that input. **Q: How well did it handle Spanish-English code-switching?** It handled mixed-language audio, but the bilingual run was much weaker than the medical sample: 28.57% WER, 57.5% Spanish token recall, and the highlighted diff dropped the Spanish token `ahora`. **Q: Does ElevenLabs Scribe return timestamps, confidence, and speaker labels?** Yes. Every scored run showed `word_timestamps yes`, `confidence yes`, and `speaker_labels yes`, with payload depth 3/3. **Q: Was speaker attribution correctness measured?** No. The benchmark measured label presence and label count, but not whether the labels were attributed to the correct speakers. **Q: What did it cost and how fast was it in this benchmark?** The run-level list-price cost came out to $0.13106, $0.06875, and $0.11857 across the three inputs, with a normalized rate of $0.2202 per audio-hour. Latency was 41.76s, 16.33s, and 6.07s, with RTFs of 0.01949, 0.01453, and 0.00313. **Q: Was streaming tested?** No. This research is a batch benchmark; it does not provide streaming-latency evidence. **Q: Does the price change across plans?** No. The report says the Scribe v2 per-hour rate stays at $0.22/hr across plans; the plans mainly change included hours and concurrency. ## Similar Tools AI tools similar to ElevenLabs: - [AssemblyAI](https://aidemos.com/tools/assemblyai-speech-to-text) — Fast batch speech-to-text with rich metadata, strong jargon and mixed-language results, but overlap-heavy meetings can still lose too much. - [Speechmatics](https://aidemos.com/tools/speechmatics) — Strong batch STT for hard English audio, but weak on code-switching as configured. - [OpenAI](https://aidemos.com/tools/openai) — Batch speech-to-text with word timestamps, but a strict upload cap and weak multilingual performance make it a mixed fit for hard audio. - [AWS Transcribe](https://aidemos.com/tools/aws-transcribe) — Batch speech-to-text with word-level metadata, but accuracy drops on overlap and code-switching. - [Rev AI](https://aidemos.com/tools/rev-ai) — Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on crosstalk and code-switching. - [Gladia](https://aidemos.com/tools/gladia) — Batch STT with rich word-level metadata and strong jargon recall, but overlap handling is weak and bilingual WER needs a mono-downmixed rerun. - [Deepgram](https://aidemos.com/tools/deepgram) — Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text transcription, audio transcription, or subtitle generation system for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).