Audio detach range without end: take the rest of a track
In Sume audio detach, range.end is optional: a start-only range runs to the end of the track. How it meets the 900 s output cap and 1800 s source cap.

In Sume audio detach, range is { start, end? } in seconds. If you leave out end, the range is open-ended and runs to the end of the track. The audio detach docs, read 2026-10-05, say this directly. Omit range altogether and you get the whole track.
Open-ended is handy when you know where the speech starts, for example after a cold open, and you want everything after it.
The caps still apply
Output is documented as capped at 900 seconds, and the source at 1800. The worker checks the 900 second cap when you send both start and end (a range over 900 seconds fails with audio_detach_range_empty, the same code used for an end before the start) and when you omit range on a source longer than 900 seconds (output_duration_exceeded). I did not find a check that rejects a start-only range whose remaining length is over 900 seconds, so do not rely on one: size the range yourself and confirm duration_seconds in the result.
| Source length | start | Remaining | Within the 900 s cap? |
|---|---|---|---|
| 1500 s | 600 | 900 s | Yes |
| 1500 s | 300 | 1200 s | No: over the cap, send an end |
| 600 s | 0 | 600 s | Yes |
Other range errors
detach_start_past_source:startis beyond the probed duration.detach_source_has_no_audio: the video has no audio stream.source_duration_exceeded: the source is over 1800 seconds.
Request
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-tail-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"range": {"start": 600},
"channels": "mono",
"sample_rate": 16000
}'The result reports duration_seconds, source_duration_seconds, and the range it used, so you can check the cut. The job costs $0.01.
Sources
Related posts
More in Media tools
- audio_source_in_requires_single_spine: source_in needs one voice file
audio.source_in only works with one audio.url. With audio.parts, set source_in on each part instead; the Sume hint is use_part_source_in.
- Avatar clip frame rate: Griffin 25 fps vs Fabric vs H3 Max
Tavus says Griffin streams 8-frame latents at 25 fps. Sume's docs say Fabric is 25 fps and H3 Max must be measured. Probe a clip with video-inspect.
- Baby shower video music: a gentle 90-second slideshow bed
Pick a gentle instrumental for a 90-second baby shower photo slideshow: one generation and a two-minute render, $0.325 on Sume.
- Put a 4:5 AI image on a 9:16 canvas with blurred fill in Pillow
Fill the empty bands of a 9:16 canvas with a blurred, darkened copy of the same Sume image, and place the sharp 4:5 original over it. Code and blur settings.
Written by Sume