Transcribe a 30-minute file: OpenAI 25 MB cap vs Sume 10-minute jobs
A 30-minute file hits OpenAI's 25 MB upload limit and Sume's 10-minute duration hint. How to split once, transcribe in pieces, and re-base the word timings.

A 30-minute recording does not fit in one request on either side. OpenAI's speech-to-text guide caps a file at 25 MB and tells you to split longer audio without breaking mid-sentence. Sume's POST /v1/stt-1.0/transcribe takes a duration_seconds hint of 1 to 600, so the working plan is three 10-minute pieces. Cut once with timeline audio, run three jobs, and add each piece's start time to its word timings.
The limits side by side
The numbers below are the ones each page states. Neither limit is a quality statement.
| Surface | Limit stated | What it means for 30 minutes |
|---|---|---|
| OpenAI speech-to-text guide | 25 MB per file; mp3, mp4, mpeg, mpga, m4a, wav, webm | Compress or split; avoid cutting mid-sentence |
| Sume STT 1.0 | duration_seconds 1 to 600; omit it and 1 minute is reserved | Three pieces of up to 10 minutes |
| Sume timeline audio split | 1 to 20 ranges per job; produced audio up to 1800 seconds | One job can cut 30 minutes into three ranges |
Step 1: get one audio file
If the recording is a video, run audio detach once to get a wav on media.sume.com. Detach takes a Sume-hosted video, so import it first with POST /v1/media-imports. A wav is sample-exact, which matters when you cut and re-time.
Step 2: split on pauses
Timeline audio split takes the file url and up to 20 ranges, each { start, end } in seconds. The result carries one audio_url per segment. Public rate is $0.01 flat per job (confirm in GET /v1/catalog).
Cutting at exactly 600 and 1200 seconds can land mid-word. Listen near each boundary, or run one quick pass that finds silences, and nudge the cut to a pause. Ranges may overlap, so you can also let pieces overlap by a second and drop duplicate words later.
Step 3: transcribe each piece, then re-base
Submit each segment's audio_url with duration_seconds: 600 and a stable Idempotency-Key. Each result returns text and words[] with times counted from the start of that piece. Add the piece's start to every word:
def merge(pieces):
"""pieces: list of (start_seconds, words) in order."""
out = []
for start, words in pieces:
for w in words:
out.append({
"word": w["word"],
"start": w["start"] + start,
"end": w["end"] + start,
})
return out
print(merge([
(0, [{"word": "Hi", "start": 0.1, "end": 0.3}]),
(600, [{"word": "again", "start": 1.0, "end": 1.4}]),
]))What it costs
At the $0.01 per audio minute list rate, 30 minutes is $0.30 of speech-to-text, plus the $0.01 split job. Microsoft's new MAI-Transcribe-2-Streaming is priced at $0.54 per hour of audio on an introductory rate through the end of 2026 (read 2026-10-04), but it is a streaming service and a different shape of integration. Poll each job as described in jobs and results rather than holding three requests open.
Sources
Related posts
More in Developers
- Transcribe audio from a private bucket: Sume STT needs a public URL
Sume STT takes a public HTTPS audio_url. If your files sit in a private bucket, here is how to hand them over without opening the whole bucket.
- Transcribe two minutes of a long video: audio detach range, then STT
Streaming transcribers charge by the hour; you may only need one segment. Detach a range as 16 kHz mono wav, then run one STT job. Caps and codes included.
- 150-language subtitles: which scripts Sume captions document
A translation model can output 150 languages, but Sume documents Latin and Hangul caption styles. Test other scripts on a short clip before a batch.
- Trigger.dev Node 21 warning: which Node runs the Sume SDK
Trigger.dev v4.6.1 added Node.js 21 deprecation warnings. The Sume TypeScript SDK needs Node 18 or later, so tasks on Node 22 or newer are fine.
Written by Sume