OpenAI transcription 25 MB limit: how many minutes of wav fit?
OpenAI caps transcription uploads at 25 MB. A 16 kHz mono wav fills that in about 13 minutes, so detach long videos to mp3 or ranges first.

A 16 kHz, mono, 16-bit wav runs 32,000 bytes per second, which is 1.92 MB per minute, so OpenAI's 25 MB transcription upload limit holds about 13 minutes of it. An mp3 at 128 kbps is 16,000 bytes per second, so the same limit holds about 26 minutes.
The 25 MB figure comes from OpenAI's File transcription guide, which says files can be up to 25 MB and lists mp3, mp4, mpeg, mpga, m4a, wav and webm as input formats. The arithmetic below is ours. If you subtitle long videos, it tells you which container to extract before you send anything.
How big is each audio format per minute?
Uncompressed PCM size is sample rate times bytes per sample times channels. Sume's audio detach page documents the two outputs you can ask for: wav (pcm_s16le) and mp3 at 128 kbps, with sample_rate of 16000, 44100 or 48000 and channels of source or mono. Those settings give the table below.
| Format | Bytes per second | MB per minute | Minutes in 25 MB |
|---|---|---|---|
| wav 16 kHz mono | 32,000 | 1.92 | about 13 |
| wav 44.1 kHz stereo | 176,400 | 10.58 | about 2.4 |
| wav 48 kHz stereo | 192,000 | 11.52 | about 2.2 |
| mp3 128 kbps | 16,000 | 0.96 | about 26 |
Why does a default detach overshoot the limit?
Detach defaults to wav and inherits the source's channels and rate. A typical phone clip is 48 kHz stereo, which is 11.52 MB per minute, so a 3-minute clip already crosses 25 MB. Ask for the STT shape instead: the Sume docs name 16000 with channels: "mono" as that shape.
Sume's own caps matter too. A detach source can be up to 1800 seconds but the output can only be 900 seconds, so a longer track needs a range. A full 900 seconds of 16 kHz mono wav is 28.8 MB, which is still over 25 MB. For a 15-minute segment sent to a 25 MB endpoint, choose mp3 (14.4 MB) or split the range in two.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-mp3-001" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "format": "mp3", "channels": "mono", "range": {"start": 0, "end": 900}}'What if you transcribe on Sume instead?
Sume's video inspect route can transcribe a Sume-hosted clip directly with transcribe: true, so there is no upload step to size. The public rate on that page is $0.01 per audio minute, and the duration hint maxes out at 600 seconds; omit it and Sume reserves one minute. Confirm the live rate in GET /v1/catalog.
The routes differ in shape. OpenAI takes an uploaded file under 25 MB. Sume takes a media.sume.com URL, so an off-host clip must be imported first with POST /v1/media-imports.
How do you pick a split for a long video?
Detach once, then cut. The Sume docs recommend this order for many ranges: run one detach for the whole track (or the first 900 seconds), then call timeline audio with operation: "split" and up to 20 ranges. Each range comes back as its own audio_url on media.sume.com.
- Keep wav when the file will be joined again; mp3 re-adds priming padding at every edge, per the timeline audio page.
- Choose ranges at sentence breaks, not at fixed 13-minute marks, or a word can be cut across two uploads.
- Transcribe each range, then add each range's start offset back to its word times before building captions.
Sources
Related posts
More in Developers
- Per-customer spend caps on Sume: what Sume caps, what you log
A Sume key belongs to a workspace, not your end customer. Cap each run with generation_spend_cap_usd and keep a per-customer ledger yourself.
- Perplexity Decisions API as a publish gate for Sume output
Check a finished Sume Format image with Perplexity's Decisions API before it ships: base64 data URL, one yes/no question, a threshold, and a human-review lane.
- Pocket TTS API: run it yourself or call a hosted TTS API
Kyutai's Pocket TTS installs with pip and serves from localhost. If you want a hosted API with job URLs instead, here is the Sume request and what changes.
- Prefect 3 task retries for a Sume job: same Idempotency-Key
Retry a Sume image job in Prefect 3 without paying twice: a tested flow with retry_condition_fn, delay list and an order-derived idempotency key.
Written by Sume