Detach 16 kHz mono audio from a 25-minute video: 2 ranges, 2 cents

A 25-minute video exceeds audio detach's 900-second output cap, so request two 750-second ranges at 16 kHz mono: two jobs at $0.01 each, about 24 MB per wav.

5 min readSume
All posts

To pull speech-to-text-ready audio out of a 25-minute Sume-hosted video, send two audio detach requests, one for seconds 0 to 750 and one for 750 to 1,500, each as 16 kHz mono wav. Detach caps output at 900 seconds, so one whole-track request fails. Two jobs at $0.01 each cost $0.02 and use only worker ffmpeg, with no provider inference.

The limits that force two jobs

Audio detach takes one video from your workspace's media.sume.com host and returns a new audio artifact. The source can be up to 1,800 seconds; the output can be up to 900 seconds. A 25-minute video is 1,500 seconds, which fits as a source but not as a single output. Use the range field with start and end in seconds. When you omit end, the range runs to the end of the track.

Detach request for a 1,500-second source (as of 2026-10-08)
Jobrange.startrange.endOutput secondsPrice
10750750$0.01
27501,500750$0.01
Total1,500$0.02

Why 16 kHz mono

The docs name sample_rate 16000 with channels mono as the speech-to-text shape. Wav output is pcm_s16le, 2 bytes per sample, so one mono channel at 16,000 samples a second is 32,000 bytes per second. A 750-second range is 24,000,000 bytes, about 24 MB. Stereo at 44.1 kHz would be about 132 MB for the same range, so the STT shape is both smaller and what the transcription step expects.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: webinar-07-part-1" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/webinar.mp4",
    "format": "wav",
    "channels": "mono",
    "sample_rate": 16000,
    "range": { "start": 0, "end": 750 }
  }'

Before you submit

Import any outside video first with POST /v1/media-imports; the server does not fetch from the open internet. Idempotency-Key is required, and each range needs its own key, as in the sample. The default mode is async; pass mode sync to wait up to 30 seconds for a completed job, otherwise poll the job status and read audio_url, duration_seconds and warnings from the result.

If the source has no audio track, the job fails with detach_source_has_no_audio. The related post on that error shows the probe that prevents it.

Next step after the two wav files

Pass each wav to speech-to-text, then stitch the transcripts using the range starts as offsets: any timestamp from the second file needs 750 seconds added. If you need the audio again later, the wav files are durable artifacts in your workspace, so there is no need to detach the same range twice.

Two jobs is the simple case. A 29-minute source is 1,740 seconds and also needs two ranges; a 14-minute source (840 seconds) fits in one job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume