Transcribe a three-hour recording when the API caps at ten minutes

Streaming sessions end at an hour; Sume STT jobs take up to 600 seconds. Split with Timeline audio, transcribe 18 chunks, and stitch word times back together.

5 min readSume
All posts

To transcribe a recording longer than your API's limit, cut it into pieces under the limit, transcribe each, and add each piece's start offset to its word times. With Sume that means POST /v1/timeline-1.0/audio with operation: split (up to 20 ranges per call) and then one POST /v1/stt-1.0/transcribe per piece, each up to 600 seconds. A three-hour file is 18 pieces of 10 minutes, which fits in one split call.

The streaming alternative has its own ceiling: Microsoft's Learn page says each MAI-Transcribe-2-Streaming session lasts up to one hour, so a three-hour call needs three sessions there too.

What are the limits you are working around?

The numbers are from Sume's OpenAPI document and Timeline audio docs, and from Microsoft Learn.

Limits read 2026-10-03 from Sume Timeline audio, the Sume STT schema and Microsoft Learn.
LimitValueWhere
STT duration_seconds1 to 600 seconds per jobSume OpenAPI
Split ranges per call1 to 20Sume Timeline audio docs
Produced audio per callUp to 1800 secondsSume Timeline audio docs
Split price$0.01 flat per jobSume Timeline audio docs
Streaming sessionUp to one hourMicrosoft Learn

How do you prepare the file?

The audio must already be on media.sume.com in your workspace; off-host URLs are rejected at admit, so import the file first with POST /v1/media-imports. If the source is a video, run audio detach once to get a wav. Timeline audio's split needs a top-level url and ranges[], and each returned segment has its own audio_url.

Two details keep the stitch honest: a split segment starts exactly at the start you asked for, so the offset to add is that same start, and duration_seconds should be the segment length so Sume reserves the right minutes.

What does the code look like?

This script splits one hosted wav into 600-second ranges (at most 20), transcribes each in turn, and prints one merged word list. Set SUME_API_KEY and SOURCE_URL, and set TOTAL_SECONDS from your file's duration.

import os, time, requests
BASE, H = "https://api.sume.com", {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
    h = dict(H, **({"Idempotency-Key": key} if key else {}))
    d = requests.post(BASE + path, headers=h, json=body, timeout=60)
    d.raise_for_status()
    d = d.json()["data"]
    while not d["terminal"]:
        time.sleep(d.get("next_poll_after_seconds") or 2)
        d = requests.get(d["status_url"], headers=H, timeout=30).json()["data"]
    r = requests.get(d["result_url"], headers=H, timeout=30)
    r.raise_for_status()  # failed or canceled jobs answer 409 here
    return r.json()["data"]["result"]

import uuid
TOTAL, STEP = int(os.environ["TOTAL_SECONDS"]), 600
starts = list(range(0, TOTAL, STEP))
assert len(starts) <= 20, "split in several calls"
ranges = [{"start": s, "end": min(s + STEP, TOTAL)} for s in starts]
parts = run("/v1/timeline-1.0/audio", {"operation": "split",
            "url": os.environ["SOURCE_URL"], "ranges": ranges},
            key="split-" + uuid.uuid4().hex)["segments"]
words = []
for seg, r in zip(parts, ranges):
    t = run("/v1/stt-1.0/transcribe", {"audio_url": seg["audio_url"],
            "duration_seconds": r["end"] - r["start"]})
    words += [dict(w, start=w["start"] + r["start"], end=w["end"] + r["start"])
              for w in t["words"]]
print(len(words), words[:2])

Why 600 seconds and not one big job?

The 600-second cap is part of the STT request schema: duration_seconds accepts 1 to 600, and a job that reserves a minute when you omit it. Splitting has side benefits: a failed chunk costs you ten minutes of work to redo instead of three hours, chunks can run in parallel up to your workspace's concurrency limit, and each chunk's text can be shown to a reviewer as soon as it finishes.

Run chunks one after another at first, as the script does, then add concurrency once you know your plan's limits. Generation admission queues accepted jobs when the workspace is at its concurrency limit, so submitting 18 at once is allowed; read Generation admission for the queue rules. If you stitch in a database, store each chunk's job id and start offset, so you can re-read results later without re-paying. Before you start, estimate the bill: 18 STT jobs at $0.01 per audio minute is $0.01 times 180 minutes, or $1.80 for three hours, plus one $0.01 split job. That figure uses the pricing code's per-minute rate and assumes you pass an accurate duration_seconds; omitting it reserves a full minute per job, which for a short final chunk reserves more than you use.

What can go wrong?

For many unrelated files rather than one long one, Batch transcription API covers the fan-out pattern. Extract 16 kHz mono audio covers the detach step.

  • A word that straddles a cut can be split or lost; overlap the ranges by a second (ranges may overlap) and drop duplicate words by time.
  • Sentence segments (segmentation.mode: sentence) restart at each chunk, so build sentences after stitching, not before.
  • Each failed chunk returns 409 from the result URL; resubmit only that chunk, never the whole file.
  • Jobs bill separately: 18 STT jobs plus one $0.01 split job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume