Transcribe two minutes of a long video: audio detach range, then STT

Streaming transcribers charge by the hour; you may only need one segment. Detach a range as 16 kHz mono wav, then run one STT job. Caps and codes included.

6 min readSume
All posts

How do you transcribe only a slice of a long video? Cut the audio first. POST /v1/audio-detach takes an optional range of { start, end } seconds and returns a wav you can hand to speech-to-text. You pay for one detach and one transcript of the slice, not the whole file.

Per-hour pricing is the headline for streaming models; Microsoft's MAI-Transcribe-2-Streaming is $0.54 per hour through year-end (read 2026-10-04). For a single meeting highlight the unit that matters is the slice.

Caps that decide the plan

These limits come from the audio detach guide and the STT contract.

Limits for detach plus STT, read 2026-10-04
LimitValue
Source video lengthAt most 1800 s
Detached audio lengthAt most 900 s; a longer whole track needs a range
STT duration_seconds hint1 to 600; omit to reserve 1 minute
STT shapeformat: wav, channels: mono, sample_rate: 16000
Detach price$0.01 per job

Two calls

The video must already be a media.sume.com artifact or asset; import first with POST /v1/media-imports. Both calls need an Idempotency-Key. The detach result carries audio_url, which is a durable Sume URL STT can read.

Set duration_seconds on the STT job to the slice length so the reservation matches what you are transcribing. Here the slice is minutes 12 to 14.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def wait(d):
    while not d["result_ready"]:
        if d.get("terminal"):
            raise RuntimeError(d.get("status"))
        time.sleep(d.get("next_poll_after_seconds") or 3)
        j = requests.get(d["status_url"], headers=H).json()
        d = j.get("data", j)
    j = requests.get(d["result_url"], headers=H).json()
    return j.get("data", j)["result"]

VIDEO = "https://media.sume.com/artifacts/artf_demo/talk.mp4"
r = requests.post(API + "/v1/audio-detach", headers={**H, "Idempotency-Key": "slice-detach-001"},
                  json={"video_url": VIDEO, "range": {"start": 720, "end": 840},
                        "channels": "mono", "sample_rate": 16000})
r.raise_for_status()
audio = wait(r.json()["data"])["audio_url"]
r = requests.post(API + "/v1/stt-1.0/transcribe", headers={**H, "Idempotency-Key": "slice-stt-001"},
                  json={"audio_url": audio, "duration_seconds": 120})
r.raise_for_status()
print(wait(r.json()["data"])["text"])

Errors you may meet

audio_detach_range_empty means end is not after start or the range is longer than 900 seconds. detach_start_past_source means the start is beyond the clip. detach_source_has_no_audio means there is no audio track; check probe.has_audio with video inspect and frames: false first. source_duration_exceeded means the source is over 1800 seconds.

Timestamps in the transcript count from the start of the slice, so add 720 seconds to map them back onto the original video.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume