Extract audio from a video for transcription: 16 kHz mono wav
Detach the audio from a Sume-hosted video as 16 kHz mono wav, then send it to speech-to-text. Costs, caps and the errors to expect.

To transcribe a video, detach its audio first as 16,000 Hz mono wav, then pass the new audio URL to speech-to-text. Sume's audio detach docs call that shape the STT shape. Detach costs $0.01 per job and leaves the video untouched. Transcription then costs $0.01 per audio minute.
Steps
- Import the video with
POST /v1/media-importsif it is not already on your workspace's media.sume.com host. Audio detach does not fetch the open internet. - Call
POST /v1/audio-detachwith anIdempotency-Keyheader andvideo_url,format: "wav",channels: "mono"andsample_rate: 16000. - Poll
GET /v1/jobs/:id/status, then readGET /v1/jobs/:id/resultforaudio_urlandduration_seconds. - Call
POST /v1/stt-1.0/transcribewith thataudio_urland the duration, rounded up to a whole second.
The limits
| Limit | Value |
|---|---|
| Detach source length | 1,800 s |
| Detach output length | 900 s; use range for longer |
| Detach price | $0.01 per job |
STT duration_seconds hint | Up to 600 s (10 minutes) |
| STT price | $0.01 per audio minute |
A long video
The duration_seconds hint on the transcribe request tops out at 600 s, so keep each request to 10 minutes of audio or less. A detach output is capped at 900 s, so for a longer video use the range field, {start, end} in seconds, with one detach per window; the source can be up to 1,800 s. If the whole track is 900 s or less, you can instead detach once and split with Timeline audio, which takes up to 20 ranges in a single $0.01 job. Add each window's start offset back to the word times when you stitch transcripts together.
Errors to expect
detach_source_has_no_audio: the video has no audio track. Check first with video inspect.audio_detach_range_empty: the range end is not after its start, or the range is over 900 s.unsupported_media_source: the URL is not on the Sume media host.
Limits
16 kHz mono is smaller and fast to process, and it discards stereo and high-frequency content; keep the original for anything you will listen to. Costs add up per window: a 30-minute video is three 10-minute windows, so three detaches ($0.03) plus 30 audio minutes of transcription ($0.30), or $0.33 in all. See also subtitle cues from sentence segments.
More in Media tools
- Facebook Reels audio spec: AAC-LC, stereo, 48 kHz, 128 kbps+
Facebook Reels wants AAC Low Complexity, stereo, 48 kHz and 128 kbps or more. What a Sume probe shows, what trim keeps, and what you must verify yourself.
- FFmpeg drawtext on a hosted API: why Sume filters refuse it, and cues
Sume video-filter refuses drawtext and subtitles because they read files. To burn text onto a clip, send cues to video-captions: text, start and end seconds.
- Adjust video brightness, contrast, saturation by API: eq filter
Send eq=brightness=0.06:contrast=1.1:saturation=1.2 as the filtergraph of Sume video-filter. Ranges come from FFmpeg: brightness -1 to 1, saturation 0 to 3.
- Fill a timeline gap with generated B-roll through an API
Premiere 26.5 can generate clips inside the timeline. Here is the API version: generate a 3 to 10 second clip, then place it in a Timeline 1.0 slot.
Written by Sume