Transcribe two minutes of a long video: audio detach range, then STT
Streaming transcribers charge by the hour; you may only need one segment. Detach a range as 16 kHz mono wav, then run one STT job. Caps and codes included.

How do you transcribe only a slice of a long video? Cut the audio first. POST /v1/audio-detach takes an optional range of { start, end } seconds and returns a wav you can hand to speech-to-text. You pay for one detach and one transcript of the slice, not the whole file.
Per-hour pricing is the headline for streaming models; Microsoft's MAI-Transcribe-2-Streaming is $0.54 per hour through year-end (read 2026-10-04). For a single meeting highlight the unit that matters is the slice.
Caps that decide the plan
These limits come from the audio detach guide and the STT contract.
| Limit | Value |
|---|---|
| Source video length | At most 1800 s |
| Detached audio length | At most 900 s; a longer whole track needs a range |
STT duration_seconds hint | 1 to 600; omit to reserve 1 minute |
| STT shape | format: wav, channels: mono, sample_rate: 16000 |
| Detach price | $0.01 per job |
Two calls
The video must already be a media.sume.com artifact or asset; import first with POST /v1/media-imports. Both calls need an Idempotency-Key. The detach result carries audio_url, which is a durable Sume URL STT can read.
Set duration_seconds on the STT job to the slice length so the reservation matches what you are transcribing. Here the slice is minutes 12 to 14.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def wait(d):
while not d["result_ready"]:
if d.get("terminal"):
raise RuntimeError(d.get("status"))
time.sleep(d.get("next_poll_after_seconds") or 3)
j = requests.get(d["status_url"], headers=H).json()
d = j.get("data", j)
j = requests.get(d["result_url"], headers=H).json()
return j.get("data", j)["result"]
VIDEO = "https://media.sume.com/artifacts/artf_demo/talk.mp4"
r = requests.post(API + "/v1/audio-detach", headers={**H, "Idempotency-Key": "slice-detach-001"},
json={"video_url": VIDEO, "range": {"start": 720, "end": 840},
"channels": "mono", "sample_rate": 16000})
r.raise_for_status()
audio = wait(r.json()["data"])["audio_url"]
r = requests.post(API + "/v1/stt-1.0/transcribe", headers={**H, "Idempotency-Key": "slice-stt-001"},
json={"audio_url": audio, "duration_seconds": 120})
r.raise_for_status()
print(wait(r.json()["data"])["text"])
Errors you may meet
audio_detach_range_empty means end is not after start or the range is longer than 900 seconds. detach_start_past_source means the start is beyond the clip. detach_source_has_no_audio means there is no audio track; check probe.has_audio with video inspect and frames: false first. source_duration_exceeded means the source is over 1800 seconds.
Timestamps in the transcript count from the start of the slice, so add 720 seconds to map them back onto the original video.
Sources
Related posts
More in Developers
- Trigger.dev Node 21 warning: which Node runs the Sume SDK
Trigger.dev v4.6.1 added Node.js 21 deprecation warnings. The Sume TypeScript SDK needs Node 18 or later, so tasks on Node 22 or newer are fine.
- Trigger.dev public tokens: keep the Sume key server-side
Trigger.dev v4.6.2 hardened authorization for public tokens. Whatever token your browser holds, a Sume API key must never be one of them. Here is the split.
- Voice replication API audit checklist before you switch
Gemini 3.8 Flash TTS is GA with voice replication and 150+ voices. Before switching providers, audit these items against Sume's live catalog.
- Reuse TTS word timestamps as caption words, skip a second STT
You already know what the voice said and when. Feed the TTS word timings to the caption job as `words` so brand names are never misheard by speech-to-text.
Written by Sume