--- title: "Rev AI" type: "AI Tool" url: "https://aidemos.com/tools/rev-ai" description: "We tested Rev AI on medical jargon, crosstalk, and bilingual code-switching: it returned structured JSON and speaker labels, but accuracy dropped hard." category: "audio-speech" website: "https://docs.rev.ai/api/asynchronous/" published: "2026-08-13T09:18:22.519045+00:00" updated: "2026-08-20T12:44:32.422241+00:00" evidenceCount: 25 verifiedCount: 21 coverage: "dense" --- # Rev AI Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on hard audio and weak multilingual recall. ## TL;DR Verdict **Structured and inexpensive, but uneven on hard audio** **Where it wins:** - You need a batch speech-to-text API that returns JSON with word timestamps, confidence values, and speaker labels. - You are okay manually auditing speaker attribution and transcript quality on difficult audio. - You want long-form audio processed asynchronously at a low per-hour cost. **Main limitation:** You need strong multilingual or code-switching accuracy on the configured English tier; Spanish recall was only 8.8%. **Pricing:** Pay As You Go — free credits $0 · Reverb Transcription $0.20 / audio-hour · Reverb Turbo Transcription $0.10 / audio-hour · Reverb Foreign Language $0.30 / audio-hour `Batch STT` · `Word timestamps` · `Speaker labels` · `3 hard audio tests` **Website:** [Visit Rev AI](https://docs.rev.ai/api/asynchronous/) ## Evidence (first-party, tested) *25 tested cells · 21/25 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:rev-ai·cross·automation-level`.* | Criterion | Scenario | Verdict | Score | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | | Automation level | cross-scenario | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/094aec4044f6433d867cc12b30b0f2b6.mov?v=1) | `ev:rev-ai·cross·automation-level` | | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 👁 observed | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·automation-level` | | Automation level | Overlapping meeting speech with cross-talk | ✓ worked | — | 👁 observed | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·automation-level` | | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 👁 observed | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·automation-level` | | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-raw-response-3-4aeb5055b3ea.json) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·export` | | Export | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-raw-response-4a8634f9a15e.json) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·export` | | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-raw-response-2-15ae5857054e.json) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·export` | | Export | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4a2ab11bafe244f8a3fc67c2645fdaf4.png?v=1) | `ev:rev-ai·cross·export` | | Export | Overlapping meeting speech and cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/fe71ec5204d8459ea0f75e4cd1e017e3.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-and-cross-talk·export` | | Export | Spanish-English code-switching conversation | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/76625b0bfceb43d4b30be38cf0b79cb9.png?v=1) | `ev:rev-ai·spanish-english-code-switching-conversation·export` | | Export | Medical anatomy dictation with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f36723578ff94150bb689c698abaeee5.png?v=1) | `ev:rev-ai·medical-anatomy-dictation-with-dense-jargon·export` | | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7abbc629d08944c5b89b2de7cb033920.png?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·input-handling` | | Input handling | Overlapping meeting speech with cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7055fd5e6fd34441a7bfe182e3ab83a3.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·input-handling` | | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/44b5d8cb148545c0808943aca942df7e.png?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·input-handling` | | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/01d57a42829b4afc9f476643ec8ea49b.jpeg?v=1) | `ev:rev-ai·cross·input-handling` | | Input handling | Medical anatomy dictation with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f36723578ff94150bb689c698abaeee5.png?v=1) | `ev:rev-ai·medical-anatomy-dictation-with-dense-jargon·input-handling` | | Input handling | Overlapping meeting speech and cross-talk | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/fe71ec5204d8459ea0f75e4cd1e017e3.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-and-cross-talk·input-handling` | | Input handling | Spanish-English code-switching conversation | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/76625b0bfceb43d4b30be38cf0b79cb9.png?v=1) | `ev:rev-ai·spanish-english-code-switching-conversation·input-handling` | | Output quality | Overlapping meeting speech with cross-talk | ⚠ struggled | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f0402ad9a65a4e3ab9d2070b19db7638.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-with-cross-talk·output-quality` | | Output quality | Bilingual Spanish-English code-switching speech | ⚠ struggled | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/dd55c61108ba4e64bb3fe74bd6cc889e.png?v=1) | `ev:rev-ai·bilingual-spanish-english-code-switching-speech·output-quality` | | Output quality | Medical anatomy narration with dense jargon | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/fc6b1c9c114e4a1bb5e46b210408f7e8.png?v=1) | `ev:rev-ai·medical-anatomy-narration-with-dense-jargon·output-quality` | | Output quality | cross-scenario | ⚠ struggled | — | 👁 observed | `ev:rev-ai·cross·output-quality` | | Output quality | Spanish-English code-switching conversation | ⚠ struggled | 25.2/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/76625b0bfceb43d4b30be38cf0b79cb9.png?v=1) | `ev:rev-ai·spanish-english-code-switching-conversation·output-quality` | | Output quality | Medical anatomy dictation with dense jargon | ◐ mixed | 9.8/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f36723578ff94150bb689c698abaeee5.png?v=1) | `ev:rev-ai·medical-anatomy-dictation-with-dense-jargon·output-quality` | | Output quality | Overlapping meeting speech and cross-talk | ⚠ struggled | 28.3/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/fe71ec5204d8459ea0f75e4cd1e017e3.png?v=1) | `ev:rev-ai·overlapping-meeting-speech-and-cross-talk·output-quality` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. ## Scored benchmark runs Three hard-audio inputs, all completed successfully. - Overlapping Speech / Crosstalk — WER 28.33%, latency 102.7s, RTF 0.04793, cost $0.11892, returned 6279 words vs 7579 reference. - Medical Jargon — WER 9.79%, jargon recall 77.8%, latency 78.32s, RTF 0.06968, cost $0.06238, returned 2786 words vs 2728 reference. - Bilingual Code-Switching — WER 25.16%, Spanish recall 8.8%, latency 99.79s, RTF 0.05148, cost $0.10759, returned 5880 words vs 6517 reference. > **Structured and inexpensive, but uneven on hard audio** > > Rev AI consistently returned structured JSON with word-level timestamps, confidence values, and speaker labels, and it was inexpensive at $0.1998 per audio-hour. The catch is accuracy: it hit 9.79% WER on medical jargon, but 28.33% on crosstalk and 25.16% on bilingual code-switching, with only 8.8% Spanish token recall. Compared with the medical-jargon clip, the same configured English tier handled overlap and code-switching much worse, so this looks like a solid batch pipeline rather than a multilingual accuracy winner unless you retest the foreign-language tier. ## Demo Recording [Video: Rev AI demo recording](https://cdn.futuresmart.ai/public/aidemos/0b24b4dda79746849bb09384eb9014ab.mov?v=1) *Video — Screen-recorded desktop walkthrough that moves from Finder into Terminal and runs the benchmark harness with Rev AI selected across the three inputs; it shows the live setup and text output, not a final polished result.* ## Feature-by-Feature Breakdown ### Asynchronous Batch Transcription Rev AI submits long-form audio as an async speech-to-text job, then lets you poll for completion and fetch the transcript when the job finishes. The member cards exercised this POST/poll/fetch flow on multiple long-audio benchmark runs, including crosstalk and other long-form cases. **Input:** > **Audio** **Output:** Transcript detail > **Image** — Transcript detail **Input:** > **Audio** **Output:** Transcript detail > **Image** — Transcript detail **Input:** > **Audio** **Output:** Transcript detail > **Image** — Transcript detail **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** The async integration pattern worked cleanly across all three files, returning transcript JSON each time; the benchmark's quality differences came from the audio difficulty, not from request handling failures. ### Structured Transcript Output Rev AI returns transcription results as structured JSON with rich metadata such as word-level timestamps, confidence values, punctuation tokens, monologues, and speaker-label fields. The member cards show this payload structure consistently across runs, supporting downstream transcript tooling. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** This structure was consistent on every input; however, the report does not separately score timestamp alignment, so the evidence here is about payload richness rather than timing accuracy. ### Speaker Diarization **Verdict:** Detected, but overlap was over-segmented Rev AI can emit speaker labels in transcript output so multi-speaker audio can be separated by speaker. The evidence includes overlapping-meeting and crosstalk samples where labels were produced, though segmentation quality varied. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** It does expose speaker labels, but the overlap case was over-segmented and the benchmark did not measure whether the labels were attributed to the right speakers. ### Code-Switching Transcription Rev AI attempts transcription on mixed-language audio, including a Spanish-English code-switching conversation. The benchmarked runs showed weak Spanish recall in the tested configuration, but both cards are exercising the same bilingual transcription capability. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** The configured English tier was effectively weak on Spanish code-switching; a fair multilingual retest would need the separate Foreign Language tier. ## Official pricing The benchmark used Reverb Transcription (English) at $0.20 per audio-hour. | Plan | Price | Notes | | --- | --- | --- | | Pay As You Go — free credits | $0 | Credits equal to 5 audio-hours of Reverb ASR; usable across all Rev AI products; no card required. | | Reverb Transcription ★ (tested) | $0.20 / audio-hour | English; this benchmark used this tier. Rounded up to the nearest second, 15-second minimum. | | Reverb Turbo Transcription | $0.10 / audio-hour | English; half price, faster, lower accuracy tier. | | Reverb Foreign Language | $0.30 / audio-hour | Spanish, French, Chinese, Portuguese + 53 more; relevant if you want a multilingual re-test. | | Whisper Fusion Transcription | $0.005 / min | English. | | Whisper Large Transcription | $0.005 / min | English. | | Human Transcription | $1.99 / min | Not an ASR product. | | Forced Alignment | $0.003 / min | Add-on. | | Language Identification | $0.003 / min | Add-on. | | Language Translation | $0.002/min standard · $0.025/min premium | Add-on. | | Summarization | $0.002/min standard · $0.025/min premium | Add-on. | | Sentiment Analysis | $0.0008 per 10 words | Add-on. | | Topic Extraction | $0.0008 per 10 words | Add-on. | | Enterprise | Volume-based | Requires sales contact; dedicated account manager, priority support, extra evaluation credits, and flexible terms. | *Free tier includes 5 audio-hours of credits. Multilingual work requires the separate Reverb Foreign Language tier; language identification and other add-ons are billed separately.* ## Is It Right For You? **Use it if** - You need a batch speech-to-text API that returns JSON with word timestamps, confidence values, and speaker labels. - You are okay manually auditing speaker attribution and transcript quality on difficult audio. - You want long-form audio processed asynchronously at a low per-hour cost. **Skip it if** - You need strong multilingual or code-switching accuracy on the configured English tier; Spanish recall was only 8.8%. - You need verified diarization correctness rather than just speaker-label presence or speaker counts. - You need streaming-latency results; this benchmark was batch-only. ## Classification - **Category:** audio-speech - **Subcategory:** audio-to-text - **Type:** speech - **Built for:** Other ## Frequently Asked Questions **Q: Did Rev AI return word-level timestamps and confidence scores?** Yes. Every raw response showed word-level tokens with `ts` and `end_ts` timestamps plus `confidence` values. **Q: Did Rev AI return speaker labels?** Yes. The raw payloads included `speaker` fields, and the benchmark detected 6 distinct speaker labels on crosstalk, 1 on medical jargon, and 3 on bilingual code-switching. **Q: How did Rev AI perform on overlapping speech?** It scored 28.33% WER on the crosstalk sample, returned 6279 words against a 7579-word reference, and over-segmented the audio into 6 speaker labels for a 4-participant file. **Q: How did Rev AI perform on medical jargon?** It scored 9.79% WER on the medical-jargon sample and achieved 77.8% jargon recall, but it still missed terms such as `cancellous` and `trabeculae`. **Q: How did Rev AI perform on bilingual code-switching?** It scored 25.16% WER on the Spanish-English sample and only 8.8% Spanish token recall (7 of 80 types), dropping or anglicising most Spanish tokens. **Q: What latency and cost did the benchmark record?** Across the three inputs, latency ranged from 78.32s to 102.75s and cost ranged from $0.06238 to $0.11892, with a list rate of $0.00333/min ($0.1998/audio-hour). **Q: Was diarization attribution correctness verified?** No. The benchmark measured speaker-label presence and counts, but it did not compute a diarization error rate or otherwise verify that each label was assigned to the right speaker. **Q: Which tier was benchmarked, and does Rev AI have a separate multilingual tier?** The benchmark used Reverb Transcription, the English tier. The pricing page lists Reverb Foreign Language separately at $0.30/audio-hour, and Language Identification is also a paid add-on. ## Similar Tools AI tools similar to Rev AI: - [Gladia](https://aidemos.com/tools/gladia) — Batch STT with rich JSON metadata and strong jargon recall, but weak overlap handling and a channel-duplication caveat on bilingual audio. - [ElevenLabs](https://aidemos.com/tools/elevenlabs) — Natural-sounding voice cloning and narration, but with only approximate voice identity. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text, audio transcription, or diarization system for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).