caption_no_speech: why a caption job fails on a silent clip
A Sume caption job returns caption_no_speech when the clip has no audible speech. The fix is cues, segments or words, not a retry.

caption_no_speech means the clip you sent has no audible speech, so speech-to-text had nothing to turn into captions. Sume's job result carries next_action: use_overlay_captions. Retrying the same request will fail the same way. To burn text onto a silent clip, send your own cues (or segments) with text, start and end, and Sume skips speech-to-text.
What triggers it and what it is not
Speech-to-captions runs when you send script_text, or when you send nothing and let speech-to-text decide the words. Both paths need audio with speech. A music-only reel, a screen recording with the mic off, or a clip whose audio track is empty fails the same way.
The docs state this is not a generic policy rejection. Your video was allowed; it simply has nothing to transcribe. That matters for how you handle it: do not rewrite the prompt or change the style, change the input form.
Four input forms and when each applies
You may send only one of script_text, words, cues and segments on a request.
| Field | Needs speech? | Use it when |
|---|---|---|
| none (speech-to-text) | Yes | The clip has a voice and you accept the transcript |
script_text | Yes | You have the script and want it aligned to the spoken words |
words | No | You have word-level times (text, start, end in seconds) |
cues / segments | No | You have phrase-level overlay text for a silent clip |
Check before you pay
Probe the clip first. A video_inspect call with frames: false returns probe.has_audio, and the inspect docs say a silent clip with transcribe: true fails with inspect_source_has_no_audio. One probe is cheaper than a failed caption job, and it tells you which input form to build.
If has_audio is true but the job still returns caption_no_speech, the audio is probably music or room tone. Send cues for that clip, or detach the audio and listen to it.
A cues request
The cue times are in seconds on the video timeline. Sume burns the text exactly as written at those times.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: overlay-cues-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/silent.mp4",
"cues": [
{"text": "Order by Friday", "start": 0.5, "end": 2.5},
{"text": "Ships in 2 days", "start": 2.5, "end": 5.0}
]
}'Where it fits in a pipeline
Treat caption_no_speech as a routing signal. A pipeline that captions mixed footage can run a free-ish probe on each clip, then send speech clips down the speech path and silent clips down the overlay path. That is cleaner than catching the error afterward, although catching it works too since next_action tells your code what to do.
Overlay text is yours to write. Sume burns exactly the text and times you send, so a silent product loop can carry a headline in the first two seconds and a call to action in the last three. Choose a Hangul style for Korean overlay text, because the Latin styles reject Hangul with a 400.
- Silent clip plus
cues: no speech-to-text, text burned at your times. - Clip with speech: do nothing special, send
video_urlalone. - Unsure: probe with
frames: falseand readprobe.has_audio.
Sources
Related posts
More in Media tools
- Captions from a video's speech: $0.20 a job, no SRT needed
Sume's video-captions endpoint transcribes the speech and burns styled captions for $0.20 a job on clips up to 60 s. Optional script_text corrects the wording.
- Check an AI voiceover by transcribing it: a script diff for 4 cents
Run Sume STT on a finished TTS take and diff it against the script to catch misread numbers and names. 4 cents per 450-character line, TTS plus the check.
- Trim and conform to 1080x1920 at 30 fps in one video-trim job
Video trim's output field re-encodes to a set width, height and 24/25/30/60 fps in the same $0.02 job, exact precision only. Request, ranges and refusal codes.
- Crop 1920x1080 to 9:16: video-filter width 0.3164, x 0.3418
To get a 9:16 strip from a 1920x1080 clip, send a Sume video-filter crop with x 0.3418, y 0, width 0.3164, height 1: about 608x1080, above TikTok's 540x960.
Written by Sume