Detach a 16 kHz mono wav from a video for speech-to-text

One call to Sume audio detach returns a 16 kHz mono wav from a video for $0.01. Cap rules, the range field for long videos, and what transcription then costs.

4 min readSume
All posts

To prepare a video for speech-to-text on Sume, call POST /v1/audio-detach with channels: "mono" and sample_rate: 16000. The docs call that combination the STT shape. It returns a new wav artifact for a flat $0.01 per job, and the video itself does not change (audio detach docs). Then send the new audio_url to POST /v1/stt-1.0/transcribe at $0.01 per audio minute.

The request

video_url must be a video on your workspace's media.sume.com. The server does not fetch from the open internet, so import other files first with POST /v1/media-imports. Idempotency-Key is required. The default mode is async, and mode: "sync" waits up to 30 seconds.

{
  "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
  "format": "wav",
  "channels": "mono",
  "sample_rate": 16000,
  "range": {"start": 0, "end": 600}
}

Caps that shape the plan

Audio detach and STT limits from the Sume docs (read 2026-10-04)
LimitValue
Source video length1,800 seconds
Detached output900 seconds, so use range beyond that
One STT job10 minutes of audio, 600 seconds
Detach price$0.01 per job
STT price$0.01 per audio minute

Plan a long recording

A 30-minute video fits the 1,800-second source cap but not the 900-second output cap or the STT cap. Detach it in three range calls of 600 seconds each, $0.03 in all, and transcribe each part. Three 10-minute STT jobs cost $0.30, so the whole video costs $0.33. If the video has no audio track the job fails with detach_source_has_no_audio, which you can predict by checking probe.has_audio with a video inspect that sets frames: false.

A shortcut, and why this still matters

Video inspect can transcribe directly with transcribe: true, at $0.01 per audio minute on top of its own compute charge (video inspect docs). Detach is the better route when you want the wav itself for other steps, such as a later join or a voice swap.

Speech models are moving to streaming. Microsoft's MAI-Transcribe-2-Streaming is priced at $0.54 per hour of audio through year end (Microsoft AI, read 2026-10-04). A prepared mono file keeps your batch path cheap whichever service reads it.

Check the output before transcribing

The detach result reports duration_seconds, format, channels and sample_rate, plus warnings[] when there are any. Read them before the STT call. If duration_seconds is longer than you planned, the range was open-ended, and sending that to STT without a duration_seconds hint reserves only one minute. Pass the real length so the reservation matches the work.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume