Transcribe a 45-minute recording: split first, offset the timestamps
A 45-minute file exceeds Sume STT's 10-minute request cap and audio detach's 1,800 s source cap. Split locally, transcribe each chunk, and stitch word times.

A 45-minute (2,700 s) recording cannot go through Sume's detach step as is: audio detach rejects sources over 1,800 s with source_duration_exceeded, and its output is capped at 900 s. STT itself takes at most 600 s per request. So pre-split the file on your machine into chunks of 10 minutes or less, host each chunk at a public HTTPS URL, transcribe each, and add the chunk's start time to every word.
Read on 2026-10-03: the Sume audio-detach and timeline-audio docs pages, the STT 1.0 request schema in the API reference, and the media-L2 and executor code they describe.
Why not detach the audio in ranges?
Detach is built for one hosted video. The docs list unsupported_media_type when the source is not a video, and the source cap is checked by the worker after probing, so ranges do not rescue a 2,700 s source. The 20-minute case (1,200 s, under the source cap) is a different story, covered in the 20-minute dubbing post linked below: there, range detaches of up to 900 s work. At 45 minutes you split before Sume.
Timeline audio can split hosted audio, but its docs say produced audio is at most 1,800 s while the worker constant I read is 900 s, so I would not plan around it for this job.
| Surface | Cap | 45-minute file |
|---|---|---|
| STT request | duration_seconds 1 to 600 | Needs 5 chunks of at most 600 s |
| Audio detach source | 1,800 s | Fails source_duration_exceeded |
| Audio detach output | 900 s | Needs a range |
| STT price | $0.01 per audio minute | About $0.45 total |
Step 1: split locally
Nine-minute chunks give five files for 45 minutes. Re-encoding to 16 kHz mono wav also gives STT a small, uniform input. Upload the chunks to storage that serves public HTTPS.
ffmpeg -i talk.m4a -ac 1 -ar 16000 -c:a pcm_s16le \
-f segment -segment_time 540 -reset_timestamps 1 chunk_%02d.wavStep 2: transcribe each chunk and offset the words
Pass the hosted chunk URLs in order. Each word's start and end are seconds from the start of its chunk, so add the chunk index times 540. Measure real chunk lengths with ffprobe if you need sample-exact offsets; a word cut at a boundary can be split across two chunks.
import os, sys, time, requests
API = "https://api.sume.com"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
CHUNK = 540
urls = sys.argv[1:]
jobs = []
for i, u in enumerate(urls):
r = requests.post(f"{API}/v1/stt-1.0/transcribe", json={
"audio_url": u, "duration_seconds": CHUNK},
headers={**AUTH, "Idempotency-Key": f"long-talk-{i:02d}"})
r.raise_for_status()
jobs.append(r.json()["data"])
words = []
for i, job in enumerate(jobs):
while not requests.get(job["status_url"], headers=AUTH).json()["data"]["terminal"]:
time.sleep(3)
res = requests.get(job["result_url"], headers=AUTH).json()["data"]["result"]
for w in res["words"]:
if w.get("type", "word") == "word":
words.append({"word": w["word"], "start": w["start"] + i * CHUNK,
"end": w["end"] + i * CHUNK})
print(len(words), words[-1])What to check after stitching
The script reserves 540 s for every chunk; the last one is shorter, so lower its duration_seconds if you want the reservation to match. Check that the final word's end is near 2,700 s. If a chunk fails, resubmit only that chunk with its own idempotency key.
Sources
Related posts
More in Developers
- Transcribe a three-hour recording when the API caps at ten minutes
Streaming sessions end at an hour; Sume STT jobs take up to 600 seconds. Split with Timeline audio, transcribe 18 chunks, and stitch word times back together.
- Transcribe an iPhone voice memo (m4a) with an API: audio_url, limits
Sume STT takes a public HTTPS audio_url up to 10 minutes. What is verified for m4a, why media-imports cannot host a memo, and a Python example.
- Translate an SRT and burn it in: Sume caption cues, limits, Python
Sume takes no SRT upload, but caption cues take the same text and times. A Python converter, the 200-cue and 60-second limits, and which fonts apply.
- Trigger.dev 4.6.3 error cause chains: log the Sume code and request id
Trigger.dev v4.6.3 shows thrown error cause chains in the dashboard, CLI and alerts. Wrap a failed Sume call so the cause carries error.code and request_id.
Written by Sume