OpenAI transcription 25 MB limit: how many minutes of wav fit?

OpenAI caps transcription uploads at 25 MB. A 16 kHz mono wav fills that in about 13 minutes, so detach long videos to mp3 or ranges first.

5 min readSume
All posts

A 16 kHz, mono, 16-bit wav runs 32,000 bytes per second, which is 1.92 MB per minute, so OpenAI's 25 MB transcription upload limit holds about 13 minutes of it. An mp3 at 128 kbps is 16,000 bytes per second, so the same limit holds about 26 minutes.

The 25 MB figure comes from OpenAI's File transcription guide, which says files can be up to 25 MB and lists mp3, mp4, mpeg, mpga, m4a, wav and webm as input formats. The arithmetic below is ours. If you subtitle long videos, it tells you which container to extract before you send anything.

How big is each audio format per minute?

Uncompressed PCM size is sample rate times bytes per sample times channels. Sume's audio detach page documents the two outputs you can ask for: wav (pcm_s16le) and mp3 at 128 kbps, with sample_rate of 16000, 44100 or 48000 and channels of source or mono. Those settings give the table below.

Audio per minute against OpenAI's 25 MB upload limit (limit read 2026-10-03, sizes computed)
FormatBytes per secondMB per minuteMinutes in 25 MB
wav 16 kHz mono32,0001.92about 13
wav 44.1 kHz stereo176,40010.58about 2.4
wav 48 kHz stereo192,00011.52about 2.2
mp3 128 kbps16,0000.96about 26

Why does a default detach overshoot the limit?

Detach defaults to wav and inherits the source's channels and rate. A typical phone clip is 48 kHz stereo, which is 11.52 MB per minute, so a 3-minute clip already crosses 25 MB. Ask for the STT shape instead: the Sume docs name 16000 with channels: "mono" as that shape.

Sume's own caps matter too. A detach source can be up to 1800 seconds but the output can only be 900 seconds, so a longer track needs a range. A full 900 seconds of 16 kHz mono wav is 28.8 MB, which is still over 25 MB. For a 15-minute segment sent to a 25 MB endpoint, choose mp3 (14.4 MB) or split the range in two.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: detach-mp3-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "format": "mp3", "channels": "mono", "range": {"start": 0, "end": 900}}'

What if you transcribe on Sume instead?

Sume's video inspect route can transcribe a Sume-hosted clip directly with transcribe: true, so there is no upload step to size. The public rate on that page is $0.01 per audio minute, and the duration hint maxes out at 600 seconds; omit it and Sume reserves one minute. Confirm the live rate in GET /v1/catalog.

The routes differ in shape. OpenAI takes an uploaded file under 25 MB. Sume takes a media.sume.com URL, so an off-host clip must be imported first with POST /v1/media-imports.

How do you pick a split for a long video?

Detach once, then cut. The Sume docs recommend this order for many ranges: run one detach for the whole track (or the first 900 seconds), then call timeline audio with operation: "split" and up to 20 ranges. Each range comes back as its own audio_url on media.sume.com.

  • Keep wav when the file will be joined again; mp3 re-adds priming padding at every edge, per the timeline audio page.
  • Choose ranges at sentence breaks, not at fixed 13-minute marks, or a word can be cut across two uploads.
  • Transcribe each range, then add each range's start offset back to its word times before building captions.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume