Check probe.has_audio before paying $0.20 for Short captions
A silent clip makes speech-based captions fail with caption_no_speech. A probe-only Video inspect with frames false shows has_audio first. Captions are $0.20.

Probe first. A Video inspect call with frames set to false returns probe facts without stills, and the docs say that is enough to read probe.has_audio. If the clip has no audio, speech-to-text captions will fail with caption_no_speech, and you should send authored cues instead. A standalone caption job is $0.20 for clips of up to 60 seconds (as of 2026-10-08), so for a 59-second Short a wrong call costs a fifth of a dollar plus a retry.
Two different failure codes
Video inspect and Video captions guard the same fact from two sides. Inspect with transcribe true on a silent clip returns inspect_source_has_no_audio. Captions without cues or words, on a silent clip, fail with caption_no_speech and a next_action of use_overlay_captions. Neither is a policy rejection.
| Call | Silent clip result | Fix |
|---|---|---|
| POST /v1/video-inspect, transcribe true | inspect_source_has_no_audio | Read probe.has_audio first |
| POST /v1/video-captions, no cues | caption_no_speech, next_action use_overlay_captions | Send cues or segments with text, start, end |
| POST /v1/video-captions with cues | Burns the authored text; no speech-to-text | None needed |
The probe call
The body is small: video_url (a media.sume.com artifact or asset in your workspace), frames false, and an Idempotency-Key. The default mode is sync, so the handler waits up to 30 seconds for a 200 and otherwise returns 202 with a job to poll. The result carries probe, and the docs name probe.has_audio as the field to read.
Inspect is billed by Modal compute, not at a fixed per-call price, and the docs give no number for a probe-only call. If you need a budget figure, read it from your usage after a trial call rather than assuming it is free.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: probe-short-001" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/short.mp4", "frames": false}'Where the Shorts rules come in
The YouTube page makes a one-minute line important: a claimed Short over 60 seconds is blocked globally. Captions help with sound-off viewing but are unrelated to claims, so do not read the $0.20 job as a fix. The reason to probe is narrower: music-only AI clips, such as a generated video with no speech, are common, and they fail the speech path.
If has_audio is true but the audio is only music, the caption job can still end in caption_no_speech, so cues are the safer path for those clips. If you already hold the script, a script_text alignment keeps the speech timings and burns your text, which is a better fit for a voiced clip than cues are.
A short decision rule
Use three branches. If has_audio is false, skip speech-to-text and send cues with text, start and end. If has_audio is true and you know the clip is voiced, send the clip with a script_text if you have the script, or let speech-to-text run without one. If has_audio is true and the track is music or effects only, treat it as the silent case. In all three the price per accepted caption job is the same $0.20, so the probe is about avoiding a wasted attempt, not about changing the rate. Remember that cues and segments are the right tool for burned-in hooks and call-to-action lines too.
Sources
Related posts
More in Developers
- Promise.allSettled for many Sume job statuses: isolate one failure
Check many Sume job statuses at once with Promise.allSettled so one 429 or network error is reported on its own row and does not hide the other results.
- Prompting Omni Flash with sound: write the audio as its own line
Gemini Omni Flash 1.1 always generates audio. A prompt layout that separates picture from sound, three example prompts, and the request on Sume.
- Python 3.15 rc3 in CI: test your Sume job poll logic before Oct 9
Python 3.15.0rc3 is out and the final is planned for Oct 9. Add a CI job that runs a stdlib unittest on your Sume status-poll decision function.
- Python argparse CLI that checks seconds per Sume model first
A 25-line argparse CLI: choices reject unknown models, a duration check uses the documented limits, and --dry-run prints the body. Runs offline.
Written by Sume