detach_source_has_no_audio: the free probe that prevents it on Sume
Sume audio detach fails with detach_source_has_no_audio on a video with no sound track. A frames-off video inspect reads probe.has_audio first. Codes inside.

detach_source_has_no_audio means the video you asked Sume to detach audio from has no audio track. Prevent it by running a video inspect with frames set to false first and reading probe.has_audio; if it is false, there is nothing to detach and you should skip the job or add audio another way.
The refusal codes, in one place
Audio detach has stable refusal codes. Most are caught at admit, before any worker runs; a few come from the worker after it probes the file. Knowing which is which tells you whether a retry can help.
| Code | Cause | Retry helps? |
|---|---|---|
| detach_source_has_no_audio | Source has no audio track (worker) | No, use another source |
| audio_detach_range_empty | range.end <= range.start, or range longer than 900 s | Yes, fix the range |
| detach_start_past_source | range.start beyond the probed duration (worker) | Yes, fix the range |
| source_duration_exceeded | Source longer than 1,800 s (worker) | No, trim or split the source first |
| unsupported_media_source | video_url not on the Sume media host | Yes, import the file |
| source_not_found | Dead media.sume.com URL, or another workspace's | Yes, use a valid URL |
| ffmpeg_fields_rejected | Client sent af, filter, ffmpeg, cmd, codec or similar | Yes, remove the fields |
The check, step by step
Video inspect with frames set to false returns the probe without stills. The inspect docs call that call sufficient for checking has_audio, and the no-transcript inspect is not billed for STT. It is the same check the transcript option needs: a silent clip fails there with inspect_source_has_no_audio.
Branch on the value. If has_audio is true, submit the detach with its own Idempotency-Key. If false, tell the user the video is silent, or send it to the voice-over path instead.
Silent AI clips are the usual cause
Some generated video models render sound and some do not, and a silent clip detaches to nothing. When you work from a mixed batch, probe every clip once and keep the result next to the asset record, so a re-run does not repeat the check.
The price of the detach itself is $0.01 per job, so the real cost of skipping the probe is the engineering time to find out why a pipeline stopped, not the cent.
Related guards
Two more details save a debugging round. The server compiles ffmpeg itself, so do not send codec or filter fields. And a whole track longer than 900 seconds needs a range, even when the source is within the 1,800-second limit.
A tiny decision table
If has_audio is true, detach. If false and the clip is meant to be spoken, the upstream generator probably produced a silent clip; regenerate it with an audio-capable model or add a voice-over with TTS and Timeline. If false and the clip is B-roll, there is nothing to do.
None of this needs a paid inference step: the probe is a read, and a detach is worker ffmpeg at $0.01 per job.
Sources
Related posts
More in Developers
- Deterministic Sume Idempotency-Key: body hash plus a take counter
Hash customer, take number and canonical body into one key. A retry replays the job; a deliberate redo bumps the take. Python, stdlib only.
- Docs page to explainer video: chunk the script under the TTS cap
Split a long document into voice-over chunks under Sume's 20,000-character TTS request limit and price them at $0.0475 per 1,000 characters. Runnable Python.
- Does a replayed sume/auto video submit get the same model and price?
Yes. The Sume docs say Auto resolves from the normalized request and catalog version, so an idempotent replay returns the same route and price.
- Does Seedance audio cost extra on Sume? generate_audio and the price
Seedance prices on Sume are per video token and carry no separate audio line. What generate_audio does, what it defaults to, and the 10 s price either way.
Written by Sume