Detach a 16 kHz mono wav from a video for speech-to-text
One call to Sume audio detach returns a 16 kHz mono wav from a video for $0.01. Cap rules, the range field for long videos, and what transcription then costs.

To prepare a video for speech-to-text on Sume, call POST /v1/audio-detach with channels: "mono" and sample_rate: 16000. The docs call that combination the STT shape. It returns a new wav artifact for a flat $0.01 per job, and the video itself does not change (audio detach docs). Then send the new audio_url to POST /v1/stt-1.0/transcribe at $0.01 per audio minute.
The request
video_url must be a video on your workspace's media.sume.com. The server does not fetch from the open internet, so import other files first with POST /v1/media-imports. Idempotency-Key is required. The default mode is async, and mode: "sync" waits up to 30 seconds.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "wav",
"channels": "mono",
"sample_rate": 16000,
"range": {"start": 0, "end": 600}
}Caps that shape the plan
| Limit | Value |
|---|---|
| Source video length | 1,800 seconds |
| Detached output | 900 seconds, so use range beyond that |
| One STT job | 10 minutes of audio, 600 seconds |
| Detach price | $0.01 per job |
| STT price | $0.01 per audio minute |
Plan a long recording
A 30-minute video fits the 1,800-second source cap but not the 900-second output cap or the STT cap. Detach it in three range calls of 600 seconds each, $0.03 in all, and transcribe each part. Three 10-minute STT jobs cost $0.30, so the whole video costs $0.33. If the video has no audio track the job fails with detach_source_has_no_audio, which you can predict by checking probe.has_audio with a video inspect that sets frames: false.
A shortcut, and why this still matters
Video inspect can transcribe directly with transcribe: true, at $0.01 per audio minute on top of its own compute charge (video inspect docs). Detach is the better route when you want the wav itself for other steps, such as a later join or a voice swap.
Speech models are moving to streaming. Microsoft's MAI-Transcribe-2-Streaming is priced at $0.54 per hour of audio through year end (Microsoft AI, read 2026-10-04). A prepared mono file keeps your batch path cheap whichever service reads it.
Check the output before transcribing
The detach result reports duration_seconds, format, channels and sample_rate, plus warnings[] when there are any. Read them before the STT call. If duration_seconds is longer than you planned, the range was open-ended, and sending that to STT without a duration_seconds hint reserves only one minute. Pass the real length so the reservation matches the work.
Sources
Related posts
More in Media tools
- Diwali gift hamper hero still: Seedream, then upscale for $0.23
A Seedream 4.5 image is $0.033 and a Sume image upscale is $0.20. One print-ready Diwali hamper hero is $0.233; four tries and one upscale is $0.33.
- Douyin ad first frame: black at most 60%, check with stills
Douyin in-feed ads need a first frame that is at most 60% black. Pull stills from your render with Sume video inspect and measure the first frame before upload.
- Edits carousels: four matching images per call
Instagram Edits now supports carousels, per secondary reports. Sume image models with an n range of 1-4 return up to four matching images per request.
- Email-sized preview image from an avatar clip: video-frames max_edge
Pull the first frame of a Sume avatar clip as a 600-pixel jpeg for an email or a link card: one video-frames call with at [0] and max_edge 600.
Written by Sume