Check has_audio first: video_inspect frames false before STT or detach
A free probe-only video_inspect tells you probe.has_audio before you reserve STT or run audio detach, so silent clips never hit the no-audio errors.

Send POST /v1/video-inspect with frames: false and Sume returns the probe alone, with no stills, so you can read probe.has_audio before you start any audio-dependent step. It matters because three Sume calls fail on a silent file: transcribe: true on inspect fails as inspect_source_has_no_audio, audio detach fails as detach_source_has_no_audio, and a standalone caption job without cues fails as caption_no_speech.
Probe and stills on inspect are unbilled, so the check costs nothing. Everything here comes from the Sume video inspect, audio detach and video captions docs.
Why check before you submit?
The failures are cheap to hit but expensive to chase in automation. An agent that runs detach, transcribe and caption in a chain will stop at the first silent clip with an error code instead of a plan. A fast probe turns that into a branch you decide in advance.
There is also a money angle on the transcript path. Speech-to-text reserves at $0.01 per audio minute, and omitting duration_seconds reserves one minute. Checking first means you never reserve against a clip with nothing to transcribe. The docs do not spell out what happens to a reservation on a failed job here, so we do not claim a refund behaviour either way; checking first avoids the question.
What does the probe call look like?
The clip must already be a media.sume.com artifact or asset in your workspace. If it lives elsewhere, import it with POST /v1/media-imports first. Idempotency-Key is required. The default mode is sync: the handler waits up to 30 seconds and answers 200 with the finished inspect, or 202 with a queued job to poll.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: probe-only-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false
}'
# read probe.has_audio in the response, then branch:
# true -> POST /v1/audio-detach or video-inspect with transcribe: true
# false -> skip, or author caption cues by handWhich step needs which gate?
The pattern: tools that need to hear something stop; tools that only place pixels warn and continue.
| Call | On a clip with no audio | Stable code or fallback |
|---|---|---|
| video-inspect with transcribe true | Fails | inspect_source_has_no_audio |
| audio-detach | Fails | detach_source_has_no_audio |
| video-captions without cues | Fails | caption_no_speech, next_action use_overlay_captions |
| video-captions with cues or segments | Burns your authored text | No speech-to-text runs |
| timeline-compose with a mute video | Succeeds with a warning | compose_video_has_no_audio |
How does this look from an agent over MCP?
The hosted MCP server exposes video_inspect as the default tool for the question what is in this clip. Writes need an idempotency_key, and under OAuth the mcp:write scope. There is no GET wrapper for inspect, so if the submit returns 202, poll with jobs_wait and then jobs_result. The MCP tools and gates page lists the tool names.
A sensible agent instruction is short: probe first with frames false, read has_audio, and only then choose between transcribe, detach or authored cues. That keeps the branch in your prompt rather than in an error handler.
What does has_audio not tell you?
It says whether an audio track exists, not whether anyone speaks on it. A clip with a music bed and no voice has audio but no speech, and a caption job without cues can still fail on it as caption_no_speech. Sume's docs describe the silent-clip case under captions as needing audible speech. If you need to know whether there is speech, run the transcript step on a short clip and look at the words; if you already know the clip is music only, go straight to cues.
How should you wire the branch into a pipeline?
Keep the probe result with the job record so a retry does not probe again. Use a stable Idempotency-Key per clip, such as the clip id plus the step name, so a repeated submit returns the original job instead of queueing a second one. A changed request under the same key is a conflict, not a new job, so revise the suffix only when the intent changes.
Then route on has_audio. With audio, either run transcribe: true on the same inspect call (the probe is free, the transcript is $0.01 per audio minute) or detach a 16 kHz mono wav with format: wav, channels: mono, sample_rate: 16000, which the audio detach docs call the speech-to-text shape. Without audio, skip the speech steps and, if you still want captions, author cues for the caption call. A silent source is also fine for timeline compose, which warns with compose_video_has_no_audio and lets the Timeline 1.0 audio spine supply sound at assemble time.
Sources
Related posts
More in Developers
- video_inspect silence_split_seconds: sentence segments for captions
How silence_split_seconds (0.2 to 3) shapes Sume video-inspect sentence segments, the 0.5 s default in the repo, and turning segments into caption cues.
- Voice API deadlines, October 2026 to February 2027
A calendar of voice and transcription API changes from vendor pages: Gemini TTS price rise, OpenAI transcription shutdown, and the xAI voice alias move.
- Which Sume audio endpoint to call: TTS, STT, music, detach, timeline
A decision map for Sume's audio API: seven endpoints, what each takes in and returns, limits and list prices, and the order they chain in.
- Sume video tools: public URL or media import first? Per tool
Video captions takes a public HTTPS URL; trim, filter, inspect, frames, compose and detach need a workspace media.sume.com clip. A tool-by-tool input guide.
Written by Sume