Transcribe a 30-minute file: OpenAI 25 MB cap vs Sume 10-minute jobs

A 30-minute file hits OpenAI's 25 MB upload limit and Sume's 10-minute duration hint. How to split once, transcribe in pieces, and re-base the word timings.

6 min readSume
All posts

A 30-minute recording does not fit in one request on either side. OpenAI's speech-to-text guide caps a file at 25 MB and tells you to split longer audio without breaking mid-sentence. Sume's POST /v1/stt-1.0/transcribe takes a duration_seconds hint of 1 to 600, so the working plan is three 10-minute pieces. Cut once with timeline audio, run three jobs, and add each piece's start time to its word timings.

The limits side by side

The numbers below are the ones each page states. Neither limit is a quality statement.

Per-request audio limits for a long recording (read 2026-10-04)
SurfaceLimit statedWhat it means for 30 minutes
OpenAI speech-to-text guide25 MB per file; mp3, mp4, mpeg, mpga, m4a, wav, webmCompress or split; avoid cutting mid-sentence
Sume STT 1.0duration_seconds 1 to 600; omit it and 1 minute is reservedThree pieces of up to 10 minutes
Sume timeline audio split1 to 20 ranges per job; produced audio up to 1800 secondsOne job can cut 30 minutes into three ranges

Step 1: get one audio file

If the recording is a video, run audio detach once to get a wav on media.sume.com. Detach takes a Sume-hosted video, so import it first with POST /v1/media-imports. A wav is sample-exact, which matters when you cut and re-time.

Step 2: split on pauses

Timeline audio split takes the file url and up to 20 ranges, each { start, end } in seconds. The result carries one audio_url per segment. Public rate is $0.01 flat per job (confirm in GET /v1/catalog).

Cutting at exactly 600 and 1200 seconds can land mid-word. Listen near each boundary, or run one quick pass that finds silences, and nudge the cut to a pause. Ranges may overlap, so you can also let pieces overlap by a second and drop duplicate words later.

Step 3: transcribe each piece, then re-base

Submit each segment's audio_url with duration_seconds: 600 and a stable Idempotency-Key. Each result returns text and words[] with times counted from the start of that piece. Add the piece's start to every word:

def merge(pieces):
    """pieces: list of (start_seconds, words) in order."""
    out = []
    for start, words in pieces:
        for w in words:
            out.append({
                "word": w["word"],
                "start": w["start"] + start,
                "end": w["end"] + start,
            })
    return out

print(merge([
    (0, [{"word": "Hi", "start": 0.1, "end": 0.3}]),
    (600, [{"word": "again", "start": 1.0, "end": 1.4}]),
]))

What it costs

At the $0.01 per audio minute list rate, 30 minutes is $0.30 of speech-to-text, plus the $0.01 split job. Microsoft's new MAI-Transcribe-2-Streaming is priced at $0.54 per hour of audio on an introductory rate through the end of 2026 (read 2026-10-04), but it is a streaming service and a different shape of integration. Poll each job as described in jobs and results rather than holding three requests open.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume