Check has_audio first: video_inspect frames false before STT or detach

A free probe-only video_inspect tells you probe.has_audio before you reserve STT or run audio detach, so silent clips never hit the no-audio errors.

5 min readSume
All posts

Send POST /v1/video-inspect with frames: false and Sume returns the probe alone, with no stills, so you can read probe.has_audio before you start any audio-dependent step. It matters because three Sume calls fail on a silent file: transcribe: true on inspect fails as inspect_source_has_no_audio, audio detach fails as detach_source_has_no_audio, and a standalone caption job without cues fails as caption_no_speech.

Probe and stills on inspect are unbilled, so the check costs nothing. Everything here comes from the Sume video inspect, audio detach and video captions docs.

Why check before you submit?

The failures are cheap to hit but expensive to chase in automation. An agent that runs detach, transcribe and caption in a chain will stop at the first silent clip with an error code instead of a plan. A fast probe turns that into a branch you decide in advance.

There is also a money angle on the transcript path. Speech-to-text reserves at $0.01 per audio minute, and omitting duration_seconds reserves one minute. Checking first means you never reserve against a clip with nothing to transcribe. The docs do not spell out what happens to a reservation on a failed job here, so we do not claim a refund behaviour either way; checking first avoids the question.

What does the probe call look like?

The clip must already be a media.sume.com artifact or asset in your workspace. If it lives elsewhere, import it with POST /v1/media-imports first. Idempotency-Key is required. The default mode is sync: the handler waits up to 30 seconds and answers 200 with the finished inspect, or 202 with a queued job to poll.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: probe-only-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": false
  }'
# read probe.has_audio in the response, then branch:
#   true  -> POST /v1/audio-detach  or  video-inspect with transcribe: true
#   false -> skip, or author caption cues by hand

Which step needs which gate?

The pattern: tools that need to hear something stop; tools that only place pixels warn and continue.

Silent-source behaviour across Sume media calls, read 2026-10-02
CallOn a clip with no audioStable code or fallback
video-inspect with transcribe trueFailsinspect_source_has_no_audio
audio-detachFailsdetach_source_has_no_audio
video-captions without cuesFailscaption_no_speech, next_action use_overlay_captions
video-captions with cues or segmentsBurns your authored textNo speech-to-text runs
timeline-compose with a mute videoSucceeds with a warningcompose_video_has_no_audio

How does this look from an agent over MCP?

The hosted MCP server exposes video_inspect as the default tool for the question what is in this clip. Writes need an idempotency_key, and under OAuth the mcp:write scope. There is no GET wrapper for inspect, so if the submit returns 202, poll with jobs_wait and then jobs_result. The MCP tools and gates page lists the tool names.

A sensible agent instruction is short: probe first with frames false, read has_audio, and only then choose between transcribe, detach or authored cues. That keeps the branch in your prompt rather than in an error handler.

What does has_audio not tell you?

It says whether an audio track exists, not whether anyone speaks on it. A clip with a music bed and no voice has audio but no speech, and a caption job without cues can still fail on it as caption_no_speech. Sume's docs describe the silent-clip case under captions as needing audible speech. If you need to know whether there is speech, run the transcript step on a short clip and look at the words; if you already know the clip is music only, go straight to cues.

How should you wire the branch into a pipeline?

Keep the probe result with the job record so a retry does not probe again. Use a stable Idempotency-Key per clip, such as the clip id plus the step name, so a repeated submit returns the original job instead of queueing a second one. A changed request under the same key is a conflict, not a new job, so revise the suffix only when the intent changes.

Then route on has_audio. With audio, either run transcribe: true on the same inspect call (the probe is free, the transcript is $0.01 per audio minute) or detach a 16 kHz mono wav with format: wav, channels: mono, sample_rate: 16000, which the audio detach docs call the speech-to-text shape. Without audio, skip the speech steps and, if you still want captions, author cues for the caption call. A silent source is also fine for timeline compose, which warns with compose_video_has_no_audio and lets the Timeline 1.0 audio spine supply sound at assemble time.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume