Transcribe a 45-minute recording: split first, offset the timestamps

A 45-minute file exceeds Sume STT's 10-minute request cap and audio detach's 1,800 s source cap. Split locally, transcribe each chunk, and stitch word times.

5 min readSume
All posts

A 45-minute (2,700 s) recording cannot go through Sume's detach step as is: audio detach rejects sources over 1,800 s with source_duration_exceeded, and its output is capped at 900 s. STT itself takes at most 600 s per request. So pre-split the file on your machine into chunks of 10 minutes or less, host each chunk at a public HTTPS URL, transcribe each, and add the chunk's start time to every word.

Read on 2026-10-03: the Sume audio-detach and timeline-audio docs pages, the STT 1.0 request schema in the API reference, and the media-L2 and executor code they describe.

Why not detach the audio in ranges?

Detach is built for one hosted video. The docs list unsupported_media_type when the source is not a video, and the source cap is checked by the worker after probing, so ranges do not rescue a 2,700 s source. The 20-minute case (1,200 s, under the source cap) is a different story, covered in the 20-minute dubbing post linked below: there, range detaches of up to 900 s work. At 45 minutes you split before Sume.

Timeline audio can split hosted audio, but its docs say produced audio is at most 1,800 s while the worker constant I read is 900 s, so I would not plan around it for this job.

Caps that decide a 45-minute plan, read 2026-10-03
SurfaceCap45-minute file
STT requestduration_seconds 1 to 600Needs 5 chunks of at most 600 s
Audio detach source1,800 sFails source_duration_exceeded
Audio detach output900 sNeeds a range
STT price$0.01 per audio minuteAbout $0.45 total

Step 1: split locally

Nine-minute chunks give five files for 45 minutes. Re-encoding to 16 kHz mono wav also gives STT a small, uniform input. Upload the chunks to storage that serves public HTTPS.

ffmpeg -i talk.m4a -ac 1 -ar 16000 -c:a pcm_s16le \
  -f segment -segment_time 540 -reset_timestamps 1 chunk_%02d.wav

Step 2: transcribe each chunk and offset the words

Pass the hosted chunk URLs in order. Each word's start and end are seconds from the start of its chunk, so add the chunk index times 540. Measure real chunk lengths with ffprobe if you need sample-exact offsets; a word cut at a boundary can be split across two chunks.

import os, sys, time, requests

API = "https://api.sume.com"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
CHUNK = 540
urls = sys.argv[1:]
jobs = []
for i, u in enumerate(urls):
    r = requests.post(f"{API}/v1/stt-1.0/transcribe", json={
        "audio_url": u, "duration_seconds": CHUNK},
        headers={**AUTH, "Idempotency-Key": f"long-talk-{i:02d}"})
    r.raise_for_status()
    jobs.append(r.json()["data"])
words = []
for i, job in enumerate(jobs):
    while not requests.get(job["status_url"], headers=AUTH).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(job["result_url"], headers=AUTH).json()["data"]["result"]
    for w in res["words"]:
        if w.get("type", "word") == "word":
            words.append({"word": w["word"], "start": w["start"] + i * CHUNK,
                          "end": w["end"] + i * CHUNK})
print(len(words), words[-1])

What to check after stitching

The script reserves 540 s for every chunk; the last one is shorter, so lower its duration_seconds if you want the reservation to match. Check that the final word's end is near 2,700 s. If a chunk fails, resubmit only that chunk with its own idempotency key.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume