Silent AI clip fails caption_no_speech: burn text with cues instead
A clip with no speech fails the Sume caption job with caption_no_speech. Send cues with text, start and end to burn authored overlay text for a Short or ad.

If you send a clip with no audible speech to Sume's caption job without any text, it fails with caption_no_speech and a next_action of use_overlay_captions. The fix is to send cues (or segments), each with text, start and end, and Sume burns your lines without speech-to-text. That is the right path for most silent AI video.
Why it happens
Speech captions use your script_text for alignment, or speech-to-text if you send none, and both need something to hear. Many generated clips are music or silence. The doc calls the error out as not a generic policy rejection, which means the clip is fine and the input mode is wrong.
The overlay request
Authored cues skip transcription. The request is the same as a speech job with the cues added.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: cues-demo-001" \
-d '{"video_url":"https://example.com/silent.mp4","style":"slam","cues":[{"text":"New: the 3-minute mug","start":0.0,"end":1.8},{"text":"Dishwasher safe","start":2.0,"end":3.5}]}'Writing good cues
Cues are your own timing, so the quality is yours. Keep each line short and give it time to be read. A line of three or four words reads comfortably in around a second and a half, and the example above uses 1.8 seconds for the first line. Leave small gaps so lines do not blur together.
Align cues to the picture. If the product appears at 2 seconds, start its line at 2.0. For a generated clip you did not time, watch it once and write the times down. A trim first can help: cut the clip to the length you want, then time the cues against the trimmed file, because the times are measured from the start of the video you send.
If the clip does have speech after all, skip the cues and send script_text so the captions follow the speaker. And if you work in Korean, the doc says Korean text on a Latin style is rejected, so pick a Hangul-capable style for those cues.
Cost and limits
Each accepted standalone caption job reserves and captures $0.20, which the doc says is for videos of up to 60 seconds under the current fixed estimate. Check the live price in the catalog. If you need captions on a longer Short, read captions on long video before you plan the batch.
Two ways the plan can go wrong. First, video_url has to be a public HTTPS URL. Second, burned-in captions are part of the picture, so viewers cannot turn them off. Keep a clean copy.
| Input | Result |
|---|---|
| Speech, no script_text | Speech-to-text, then burn |
| Speech plus script_text | Align your text to the speech |
| Silent clip, no cues | Fails with caption_no_speech |
| Silent clip plus cues | Burns authored overlay text |
Sources
Related posts
More in Media tools
- Silent menu-special clip: authored captions on 12 seconds for $0.22
Caption a silent 12-second menu-special clip with authored cues instead of speech: $0.20 for captions plus $0.02 to trim it first, $0.22 on Sume.
- 6-second ad legal line: Sume TTS speed 1.5 and checking word times
A fast legal read tops out at generation_config.speed 1.5 on Sume TTS. Read the word timings to see whether the line fits your 6 seconds, then cut words.
- Slide-up transitions for a vertical Short: Sume Timeline types, limits
Sume Timeline has six transitions, including slideup and slidedown that suit a vertical scroll. Each is at most 1 second and half the shorter clip.
- Slow breathing voiceover: Sume TTS speed 0.6 is the floor
For a calm, slow voiceover set generation_config.speed as low as 0.6 on Sume TTS. Below that the request is rejected, so slow the rest through writing.
Written by Sume