--- title: "Rev AI" type: "AI Tool" url: "https://aidemos.com/tools/rev-ai" description: "On our batch audio test, Rev AI returned structured JSON with word timestamps, confidences, and speaker labels—but failed on crosstalk and code-switching." category: "audio-speech" website: "https://docs.rev.ai/api/asynchronous/" published: "2026-08-13T09:18:22.519045+00:00" updated: "2026-09-01T02:59:51.394002+00:00" evidenceCount: 16 verifiedCount: 16 coverage: "dense" --- # Rev AI Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on crosstalk and code-switching. ## TL;DR Verdict **Good value for batch transcription, but not a safe default for hard multilingual audio.** **Where it wins:** - You need a low-cost batch STT API that returns structured JSON with word-level timestamps, confidence values, and speaker labels. - You can process long audio asynchronously and are fine auditing difficult clips manually. - You want measured throughput around $0.1998/audio-hour rather than a higher-cost transcription service. **Main limitation:** You need verified speaker attribution on overlap; the crosstalk run over-segmented speakers and DER was not measured. **Pricing:** Pay As You Go — free credits $0 · Reverb Transcription $0.20 / audio-hour · Reverb Turbo Transcription $0.10 / audio-hour · Reverb Foreign Language $0.30 / audio-hour `Batch STT` · `Word timestamps` · `Speaker labels` · `8.8% Spanish recall` **Website:** [Visit Rev AI](https://docs.rev.ai/api/asynchronous/) ## Evidence (first-party, tested) *16 tested cells · 16/16 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:rev-ai·cross·automation-level`.* | Criterion | Scenario | Verdict | Score | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | | Automation level | cross-scenario | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/094aec4044f6433d867cc12b30b0f2b6.mov?v=1) | `ev:rev-ai·cross·automation-level` | | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a8441d00948d4bcfb4c569ee592ed81e.png?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·automation-level` | | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5ae09ce9f486423cbca1fb334cd66af7.png?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·automation-level` | | Automation level | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/3f3f9d86a4e24b28b6e8b5c05c51e846.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·automation-level` | | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/6ce27357d73c4f53bde1e1d6610b0ea4.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·export` | | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7822d559a8854a67a800a2c543183912.png?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·export` | | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/d155259eeb5f417a963a46cf3b7339fb.png?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·export` | | Export | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/rev-ai-rev-ai-57b888ecc885.md) | `ev:rev-ai·cross·export` | | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/50f3d66e58ca4d0abecc6b021fa45e62.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·input-handling` | | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-rev-ai-57b888ecc885.md) | `ev:rev-ai·cross·input-handling` | | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/012cfa7be3e94f558e8b420dd26cb8a1.png?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·input-handling` | | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a777b5de6fea4b6693c6e9ba2e5cde69.png?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·input-handling` | | Output quality | cross-scenario | ◐ mixed | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-rev-ai-57b888ecc885.md) | `ev:rev-ai·cross·output-quality` | | Output quality | Overlapping meeting speech with cross-talk | ⚠ struggled | 28.3/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/ca7ec1d4e4744cd4b2b97e681b5fce6f.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·output-quality` | | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 9.8/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/a3e7f32466ec42768390a407f03d7174.png?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·output-quality` | | Output quality | Bilingual Spanish-English code-switching speech | ⚠ struggled | 25.2/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7b4c69f4fd9140dd9bd5aeab288fdd7e.png?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·output-quality` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. > **Good value for batch transcription, but not a safe default for hard multilingual audio.** > > Rev AI consistently returned structured transcript JSON with word-level timestamps, confidence values, and speaker labels, and the benchmarked Reverb Transcription tier came out to $0.1998/audio-hour. The tradeoff is accuracy: it scored 9.79% WER on medical jargon, but 28.33% on overlapping speech and 25.16% on bilingual code-switching, with only 8.8% Spanish token recall. ## Demo Recording [Video: Rev AI demo recording](https://cdn.futuresmart.ai/public/aidemos/0b24b4dda79746849bb09384eb9014ab.mov?v=1) *Video — Desktop walkthrough of the benchmark setup in Finder and Terminal; it shows Rev AI selected for the three benchmark inputs and the live benchmark output, not a final polished result.* ## Feature-by-Feature Breakdown ### Asynchronous Batch Transcription Rev AI submits long-form audio as asynchronous speech-to-text jobs, then lets you poll and fetch the transcript once processing completes. The cards exercise this POST → poll → fetch flow on multiple long-audio files, including crosstalk.wav, medical_terms.mp3, and mix_language.mp3. **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Bottom line:** The async integration pattern worked cleanly across all three files, returning transcript JSON each time; the benchmark's quality differences came from the audio difficulty, not from request handling failures. ### Structured Transcript Output with Metadata **Verdict:** The API consistently exposed rich transcript metadata for downstream parsing. Rev AI returns transcript JSON with rich metadata such as word-level timestamps, confidence values, punctuation, timed word tokens, monologues, and speaker-label fields. The cards exercise this structured payload on repeated runs and a word-level metadata variant. **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Bottom line:** The structured payload is consistent and integration-friendly across all three runs. ### Speaker Diarization **Verdict:** Speaker labels are always present, but diarization correctness was not verified. Rev AI emits speaker labels in transcript output so multi-speaker audio can be separated by speaker. The cards cover crosstalk and overlapping-meeting samples where labels were present, though segmentation quality varied. **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Bottom line:** Speaker-label presence is reliable, but crosstalk was over-segmented and attribution quality remains unverified. ### Multilingual and Code-Switching Transcription **Verdict:** The benchmarked English tier handled Spanish-English code-switching poorly. Rev AI attempts transcription on mixed-language audio, including English/Spanish code-switching examples. The cards show the capability on bilingual samples, but also note weak Spanish recall in the tested configuration. **Input:** > **Audio** **Output:** **Bottom line:** Do not treat the benchmarked English configuration as a safe multilingual default; Spanish-heavy code-switching was mostly mistranscribed. ## Official pricing The benchmarked Reverb Transcription tier is $0.20/audio-hour, and Rev AI also offers a free-credit entry tier plus a separate multilingual tier. | Plan | Price | Notes | | --- | --- | --- | | Pay As You Go — free credits | $0 | Credits equal to 5 audio-hours of Reverb ASR; usable across all Rev AI products; no card required. | | Reverb Transcription ★ (tested) | $0.20 / audio-hour | English; benchmarked tier; rounded up to the nearest second; 15-second minimum. | | Reverb Turbo Transcription | $0.10 / audio-hour | English; lower-accuracy tier. | | Reverb Foreign Language | $0.30 / audio-hour | Spanish, French, Chinese, Portuguese + 53 more; separate multilingual tier. | *Source: https://www.rev.ai/pricing, accessed 2026-08-14. The benchmark used Reverb Transcription.* ## Is It Right For You? **Use it if** - You need a low-cost batch STT API that returns structured JSON with word-level timestamps, confidence values, and speaker labels. - You can process long audio asynchronously and are fine auditing difficult clips manually. - You want measured throughput around $0.1998/audio-hour rather than a higher-cost transcription service. **Skip it if** - You need verified speaker attribution on overlap; the crosstalk run over-segmented speakers and DER was not measured. - You need strong Spanish or other code-switching performance on the benchmarked English tier; Spanish recall was only 8.8%. - You need streaming latency results; this benchmark was batch-only. ## Classification - **Category:** audio-speech - **Subcategory:** audio-to-text - **Type:** speech - **Built for:** Other ## Frequently Asked Questions **Q: Did Rev AI return word-level timestamps and confidence scores?** Yes. Every run showed word_timestamps and confidence metadata in the raw response, along with timed word tokens. **Q: Did Rev AI return speaker labels?** Yes. Speaker_labels were detected on all three runs. The crosstalk clip showed 6 labels, the medical-jargon clip 1, and the bilingual clip 3, but attribution correctness was not scored. **Q: How did Rev AI perform on overlapping speech?** It scored 28.33% WER on the crosstalk sample, with 601 substitutions, 1423 deletions, and 123 insertions. The transcript detail panel also showed the meeting over-segmented into 6 speaker labels for a 4-participant call. **Q: How did Rev AI perform on medical jargon?** Better than the hard audio cases: 9.79% WER and 77.8% jargon recall. The report still noted misses such as cancellous and trabeculae. **Q: How did Rev AI perform on bilingual code-switching?** Poorly on the benchmarked English tier: 25.16% WER and only 8.8% Spanish token recall. The report says a fair multilingual retest would need the separate Reverb Foreign Language tier. **Q: What did Rev AI cost in this benchmark?** The benchmarked Reverb Transcription tier was priced at $0.20/audio-hour, which the report also records as $0.00333/min. Measured run costs were $0.11892, $0.06238, and $0.10759 across the three inputs. **Q: Is there a free tier?** Yes. The pricing section lists Pay As You Go free credits equal to 5 audio-hours of Reverb ASR, with no card required. **Q: Was diarization attribution correctness verified?** No. The benchmark detected speaker labels and counted them, but it did not compute DER or otherwise score whether the labels were assigned to the right speakers. **Q: Was streaming latency measured?** No. This was a batch benchmark, so latency was measured as wall-clock time and real-time factor rather than live streaming lag. ## Similar Tools AI tools similar to Rev AI: - [AssemblyAI](https://aidemos.com/tools/assemblyai-speech-to-text) — Fast batch speech-to-text with rich metadata, strong jargon and mixed-language results, but overlap-heavy meetings can still lose too much. - [Speechmatics](https://aidemos.com/tools/speechmatics) — Strong batch STT for hard English audio, but weak on code-switching as configured. - [OpenAI](https://aidemos.com/tools/openai) — Batch speech-to-text with word timestamps, but a strict upload cap and weak multilingual performance make it a mixed fit for hard audio. - [AWS Transcribe](https://aidemos.com/tools/aws-transcribe) — Batch speech-to-text with word-level metadata, but accuracy drops on overlap and code-switching. - [ElevenLabs Scribe](https://aidemos.com/tools/elevenlabs-scribe) — Fast batch speech-to-text with word-level metadata, strongest on jargon and weaker on overlap/code-switching. - [Gladia](https://aidemos.com/tools/gladia) — Batch STT with rich word-level metadata and strong jargon recall, but overlap handling is weak and bilingual WER needs a mono-downmixed rerun. - [Deepgram](https://aidemos.com/tools/deepgram) — Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching. - [ElevenLabs](https://aidemos.com/tools/elevenlabs) — Natural-sounding voice cloning and narration, but with only approximate voice identity. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text, audio transcription, or transcription system for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).