--- title: "OpenAI" type: "AI Tool" url: "https://aidemos.com/tools/openai-speech-to-text" description: "We sent OpenAI Speech-to-Text a jargon-heavy clip and got usable word-level JSON timestamps, but the 65.39 MB crosstalk file hit its 25 MB cap." category: "audio-speech" website: "https://platform.openai.com/docs/guides/speech-to-text" published: "2026-08-27T12:46:18.793206+00:00" updated: "2026-09-01T02:55:25.640203+00:00" evidenceCount: 16 verifiedCount: 16 coverage: "dense" --- # OpenAI Batch speech-to-text with word timestamps, but a strict 25 MB cap and weak code-switching make it a mixed fit for hard audio. ## TL;DR Verdict **Useful on small, mostly-English audio; mixed on hard cases.** **Where it wins:** - You need a simple hosted batch STT API that returns word-level timestamps. - Your audio files are under 25 MB. - Your audio is mostly English or jargon-heavy narration rather than strongly code-switched speech. **Main limitation:** You need speaker labels or confidence scores. **Pricing:** whisper-1 $0.006/min · gpt-4o-transcribe $0.006/min · gpt-4o-mini-transcribe $0.003/min · gpt-4o-transcribe-diarize $0.006/min `25 MB cap` · `Word timestamps` · `Jargon tested` · `Code-switching weak` **Website:** [Visit OpenAI](https://platform.openai.com/docs/guides/speech-to-text) ## Evidence (first-party, tested) *16 tested cells · 16/16 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:openai·cross·automation-level`.* | Criterion | Scenario | Verdict | Score | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | | Automation level | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/779f9a413ae14ff5b8984ac4c7debb5e.png?v=1) | `ev:openai·cross·automation-level` | | Automation level | Overlapping meeting speech with cross-talk | ✗ failed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/96638b98259c4d2aac95d1f588cca038.png?v=1) | `ev:openai·overlapping-meeting-speech-with-cross-talk·automation-level` | | Automation level | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4446fa2a39d844bab21fd46987159402.png?v=1) | `ev:openai·medical-anatomy-narration-with-dense-jargon·automation-level` | | Automation level | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/ed1e8c921bcd4ced8e5aa03d04b01b26.png?v=1) | `ev:openai·bilingual-spanish-english-code-switching-speech·automation-level` | | Export | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4aa7c6faa826460ba29d6b929ec042ee.png?v=1) | `ev:openai·bilingual-spanish-english-code-switching-speech·export` | | Export | Overlapping meeting speech with cross-talk | ✗ failed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/354b416409de410091540e2fff6d7877.png?v=1) | `ev:openai·overlapping-meeting-speech-with-cross-talk·export` | | Export | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/04bcddb387974e7693066d4bbcefed6d.png?v=1) | `ev:openai·medical-anatomy-narration-with-dense-jargon·export` | | Export | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/5e38d0613a324b269256234bda0e1c29.png?v=1) | `ev:openai·cross·export` | | Input handling | Bilingual Spanish-English code-switching speech | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f61c6be7a62a472dbf02d38312f178b7.png?v=1) | `ev:openai·bilingual-spanish-english-code-switching-speech·input-handling` | | Input handling | Medical anatomy narration with dense jargon | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4446fa2a39d844bab21fd46987159402.png?v=1) | `ev:openai·medical-anatomy-narration-with-dense-jargon·input-handling` | | Input handling | cross-scenario | ✓ worked | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/7ad7d217f66e4740a2344a4cf082d94d.jpeg?v=1) | `ev:openai·cross·input-handling` | | Input handling | Overlapping meeting speech with cross-talk | ✗ failed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/2d9859ef01c9489e848ac12a440168a8.png?v=1) | `ev:openai·overlapping-meeting-speech-with-cross-talk·input-handling` | | Output quality | cross-scenario | ◐ mixed | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/research-media-openai-whisper-66c2d3e41cd0.md) | `ev:openai·cross·output-quality` | | Output quality | Overlapping meeting speech with cross-talk | ◐ mixed | — | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/2d9859ef01c9489e848ac12a440168a8.png?v=1) | `ev:openai·overlapping-meeting-speech-with-cross-talk·output-quality` | | Output quality | Medical anatomy narration with dense jargon | ✓ worked | 5.2/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/4446fa2a39d844bab21fd46987159402.png?v=1) | `ev:openai·medical-anatomy-narration-with-dense-jargon·output-quality` | | Output quality | Bilingual Spanish-English code-switching speech | ⚠ struggled | 26.1/5 | 🧾 [proof](https://cdn.futuresmart.ai/public/aidemos/f61c6be7a62a472dbf02d38312f178b7.png?v=1) | `ev:openai·bilingual-spanish-english-code-switching-speech·output-quality` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. ## Across the three inputs Best on jargon, weak on code-switching, and hard-limited on upload size. - **Best** Medical jargon — 5.17% WER and 88.9% jargon recall on the anatomy sample. - **Weak** Bilingual code-switching — 26.12% WER and 46.2% Spanish token recall on a mostly-English clip. - **Blocked** Overlapping speech / crosstalk — Rejected with HTTP 413 because the 65.39 MB file exceeded the 25 MB cap. > **Useful on small, mostly-English audio; mixed on hard cases.** > > Benchmarked on whisper-1, OpenAI Speech-to-Text returned usable transcripts on the jargon-heavy sample and exposed word-level timestamps in verbose JSON. But the 65.39 MB crosstalk file was hard-rejected by the API's 25 MB cap, and the bilingual clip only recovered 46.2% of Spanish tokens, so this is a practical batch API for smaller English-heavy audio rather than a strong multilingual or meeting-style engine. ## Demo Recording [Video: OpenAI demo recording](https://cdn.futuresmart.ai/public/aidemos/d019dd557ac74e6eb7fe20c7ed0cf19d.mov?v=1) *Video — Screen recording of the benchmark workflow and result views for OpenAI Speech-to-Text.* ## Feature-by-Feature Breakdown ### Audio Transcription **Verdict:** Works on sub-25 MB files, but hard-rejects oversized audio. Transcribes uploaded audio into text, including multipart uploads under the documented size cap. The exercised inputs included jargon-heavy English medical narration, a Spanish-English mixed clip, and larger multipart files such as the 8.6 MB medical narration and 22.2 MB bilingual clip. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Good for small batch uploads; the 25 MB cap is a real deployment constraint. ### Word-Level Timestamped Transcript Output **Verdict:** Timed tokens are present, but there is no confidence or speaker-label signal. Returns transcript output with word-level timing metadata when verbose JSON and timestamp granularity are enabled. The exercised outputs were structured JSON exports with timed tokens for captioning, search alignment, and downstream synchronization. **Input:** > **Audio** **Output:** > **Image** **Input:** > **Audio** **Output:** > **Image** **Bottom line:** Useful alignment metadata, but not enough for diarization or confidence gating. ## Official list prices from OpenAI's pricing page Per-minute API pricing; whisper-1 is the benchmarked model on this page. | Plan | Price | Notes | | --- | --- | --- | | whisper-1 ★ (tested) | $0.006/min | Flat per-minute pricing; this was the benchmarked model. | | gpt-4o-transcribe | $0.006/min | Vendor lists this as an estimated cost column. | | gpt-4o-mini-transcribe | $0.003/min | Cheapest 4o-family transcribe model listed on the pricing page. | | gpt-4o-transcribe-diarize | $0.006/min | Diarization is included at the same rate; it is not a separate surcharge. | | gpt-transcribe | $0.0045/min | Newer non-4o transcribe model. | | gpt-live-transcribe | $0.017/min | Live/streaming transcription. | *Prices were read from OpenAI's pricing page on 2026-08-14 and should be re-verified at test time. No free tier or free credits are published on the API pricing page, and Batch API is not priced separately for transcription.* ## Is It Right For You? **Use it if** - You need a simple hosted batch STT API that returns word-level timestamps. - Your audio files are under 25 MB. - Your audio is mostly English or jargon-heavy narration rather than strongly code-switched speech. **Skip it if** - You need speaker labels or confidence scores. - You need balanced multilingual transcription or strong code-switching support. - Your audio routinely exceeds 25 MB or includes long meeting recordings over the cap. ## Classification - **Category:** audio-speech - **Subcategory:** audio-to-text - **Type:** text - **Built for:** Other ## Frequently Asked Questions **Q: What file-size limit did OpenAI hit in this benchmark?** A 65.39 MB WAV was rejected with HTTP 413 because the documented upload cap is 25 MB. The failed request took 50.08 seconds and was charged $0.0. **Q: Does it return word-level timestamps?** Yes. With verbose_json and timestamp_granularities[] word enabled, the accepted runs returned timed tokens; the raw response shows 2,664 word-level tokens on the medical sample and 5,703 on the bilingual sample. **Q: Does it include speaker labels or confidence scores?** No. The raw response summary shows confidence=no and speaker_labels=no on both accepted runs. **Q: How accurate was it on the medical jargon sample?** WER was 5.17% on the anatomy sample, with 88.9% jargon recall. The highlighted miss was 'trabeculae'. **Q: How did it perform on the bilingual code-switching sample?** WER was 26.12%, and Spanish token recall was 46.2% (37/80 types) on a clip that was mostly English, so balanced multilingual performance is not demonstrated here. **Q: What does it cost and is there a free tier?** OpenAI's pricing page lists whisper-1 at $0.006/min ($0.36/hour). The report also lists gpt-4o-mini-transcribe at $0.003/min and gpt-transcribe at $0.0045/min. No free tier or free credits are published on the API pricing page. **Q: What did it cost on the benchmarked runs?** The medical-jargon run cost $0.1124 and the bilingual run cost $0.19385 at the report's list-price assumptions. The oversize overlapping-speech run was charged $0.0 because it was rejected before transcription. ## Similar Tools AI tools similar to OpenAI: - [Gladia](https://aidemos.com/tools/gladia) — Batch STT with rich JSON metadata and strong jargon recall, but weak overlap handling and a channel-duplication caveat on bilingual audio. - [AssemblyAI](https://aidemos.com/tools/assemblyai-speech-to-text) — Fast batch STT with strong metadata and mixed-language performance, but overlap-heavy meetings can drop too many words. - [Speechmatics](https://aidemos.com/tools/speechmatics) — Strong batch STT for hard English audio, but weak on code-switching as configured. - [Google Cloud Speech-to-Text](https://aidemos.com/tools/google-cloud-speech-to-text) — Batch transcription with word timestamps, but uneven accuracy on overlap and code-switching. - [AWS Transcribe](https://aidemos.com/tools/aws-transcribe) — Batch speech-to-text with timestamps and speaker labels, but weak on multilingual audio. - [ElevenLabs Scribe](https://aidemos.com/tools/elevenlabs-scribe) — Fast batch speech-to-text with word-level metadata, strongest on jargon and weaker on overlap/code-switching. - [Rev AI](https://aidemos.com/tools/rev-ai) — Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on hard audio and weak multilingual recall. - [Deepgram](https://aidemos.com/tools/deepgram) — Batch speech-to-text with rich metadata, but weak on crosstalk and code-switching. - [ElevenLabs](https://aidemos.com/tools/elevenlabs) — Natural-sounding voice cloning and narration, but with only approximate voice identity. ## Need a custom AI solution for this use case? If you are looking to build a custom speech-to-text, audio transcription, or word timestamping system for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).