Transcribe a three-hour recording when the API caps at ten minutes
Streaming sessions end at an hour; Sume STT jobs take up to 600 seconds. Split with Timeline audio, transcribe 18 chunks, and stitch word times back together.

To transcribe a recording longer than your API's limit, cut it into pieces under the limit, transcribe each, and add each piece's start offset to its word times. With Sume that means POST /v1/timeline-1.0/audio with operation: split (up to 20 ranges per call) and then one POST /v1/stt-1.0/transcribe per piece, each up to 600 seconds. A three-hour file is 18 pieces of 10 minutes, which fits in one split call.
The streaming alternative has its own ceiling: Microsoft's Learn page says each MAI-Transcribe-2-Streaming session lasts up to one hour, so a three-hour call needs three sessions there too.
What are the limits you are working around?
The numbers are from Sume's OpenAPI document and Timeline audio docs, and from Microsoft Learn.
| Limit | Value | Where |
|---|---|---|
STT duration_seconds | 1 to 600 seconds per job | Sume OpenAPI |
| Split ranges per call | 1 to 20 | Sume Timeline audio docs |
| Produced audio per call | Up to 1800 seconds | Sume Timeline audio docs |
| Split price | $0.01 flat per job | Sume Timeline audio docs |
| Streaming session | Up to one hour | Microsoft Learn |
How do you prepare the file?
The audio must already be on media.sume.com in your workspace; off-host URLs are rejected at admit, so import the file first with POST /v1/media-imports. If the source is a video, run audio detach once to get a wav. Timeline audio's split needs a top-level url and ranges[], and each returned segment has its own audio_url.
Two details keep the stitch honest: a split segment starts exactly at the start you asked for, so the offset to add is that same start, and duration_seconds should be the segment length so Sume reserves the right minutes.
What does the code look like?
This script splits one hosted wav into 600-second ranges (at most 20), transcribes each in turn, and prints one merged word list. Set SUME_API_KEY and SOURCE_URL, and set TOTAL_SECONDS from your file's duration.
import os, time, requests
BASE, H = "https://api.sume.com", {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
h = dict(H, **({"Idempotency-Key": key} if key else {}))
d = requests.post(BASE + path, headers=h, json=body, timeout=60)
d.raise_for_status()
d = d.json()["data"]
while not d["terminal"]:
time.sleep(d.get("next_poll_after_seconds") or 2)
d = requests.get(d["status_url"], headers=H, timeout=30).json()["data"]
r = requests.get(d["result_url"], headers=H, timeout=30)
r.raise_for_status() # failed or canceled jobs answer 409 here
return r.json()["data"]["result"]
import uuid
TOTAL, STEP = int(os.environ["TOTAL_SECONDS"]), 600
starts = list(range(0, TOTAL, STEP))
assert len(starts) <= 20, "split in several calls"
ranges = [{"start": s, "end": min(s + STEP, TOTAL)} for s in starts]
parts = run("/v1/timeline-1.0/audio", {"operation": "split",
"url": os.environ["SOURCE_URL"], "ranges": ranges},
key="split-" + uuid.uuid4().hex)["segments"]
words = []
for seg, r in zip(parts, ranges):
t = run("/v1/stt-1.0/transcribe", {"audio_url": seg["audio_url"],
"duration_seconds": r["end"] - r["start"]})
words += [dict(w, start=w["start"] + r["start"], end=w["end"] + r["start"])
for w in t["words"]]
print(len(words), words[:2])
Why 600 seconds and not one big job?
The 600-second cap is part of the STT request schema: duration_seconds accepts 1 to 600, and a job that reserves a minute when you omit it. Splitting has side benefits: a failed chunk costs you ten minutes of work to redo instead of three hours, chunks can run in parallel up to your workspace's concurrency limit, and each chunk's text can be shown to a reviewer as soon as it finishes.
Run chunks one after another at first, as the script does, then add concurrency once you know your plan's limits. Generation admission queues accepted jobs when the workspace is at its concurrency limit, so submitting 18 at once is allowed; read Generation admission for the queue rules. If you stitch in a database, store each chunk's job id and start offset, so you can re-read results later without re-paying. Before you start, estimate the bill: 18 STT jobs at $0.01 per audio minute is $0.01 times 180 minutes, or $1.80 for three hours, plus one $0.01 split job. That figure uses the pricing code's per-minute rate and assumes you pass an accurate duration_seconds; omitting it reserves a full minute per job, which for a short final chunk reserves more than you use.
What can go wrong?
For many unrelated files rather than one long one, Batch transcription API covers the fan-out pattern. Extract 16 kHz mono audio covers the detach step.
- A word that straddles a cut can be split or lost; overlap the ranges by a second (ranges may overlap) and drop duplicate words by time.
- Sentence segments (
segmentation.mode: sentence) restart at each chunk, so build sentences after stitching, not before. - Each failed chunk returns 409 from the result URL; resubmit only that chunk, never the whole file.
- Jobs bill separately: 18 STT jobs plus one $0.01 split job.
Sources
Related posts
More in Developers
- Translate an SRT and burn it in: Sume caption cues, limits, Python
Sume takes no SRT upload, but caption cues take the same text and times. A Python converter, the 200-cue and 60-second limits, and which fonts apply.
- Validate video duration and resolution in Python before you submit
Fetch GET /v1/videos/models and check duration, resolution and aspect_ratio per model in about 25 lines of Python, before a Sume video job fails.
- AI video API fallback: retry on another model when a job fails
Chain seedance-2.5, seedance-2 and kling-3 on Sume: poll status_url, read the job error category, and resubmit the brief to the next model.
- Virtual try-on API: which Sume call returns an image, which a video
Need a try-on photo or a try-on clip? On Sume the two catalog try-on Formats return video; a still comes from the image API. The table, plus one call for each.
Written by Sume