Silent AI clip fails caption_no_speech: burn text with cues instead

A clip with no speech fails the Sume caption job with caption_no_speech. Send cues with text, start and end to burn authored overlay text for a Short or ad.

4 min readSume
All posts

If you send a clip with no audible speech to Sume's caption job without any text, it fails with caption_no_speech and a next_action of use_overlay_captions. The fix is to send cues (or segments), each with text, start and end, and Sume burns your lines without speech-to-text. That is the right path for most silent AI video.

Why it happens

Speech captions use your script_text for alignment, or speech-to-text if you send none, and both need something to hear. Many generated clips are music or silence. The doc calls the error out as not a generic policy rejection, which means the clip is fine and the input mode is wrong.

The overlay request

Authored cues skip transcription. The request is the same as a speech job with the cues added.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: cues-demo-001" \
  -d '{"video_url":"https://example.com/silent.mp4","style":"slam","cues":[{"text":"New: the 3-minute mug","start":0.0,"end":1.8},{"text":"Dishwasher safe","start":2.0,"end":3.5}]}'

Writing good cues

Cues are your own timing, so the quality is yours. Keep each line short and give it time to be read. A line of three or four words reads comfortably in around a second and a half, and the example above uses 1.8 seconds for the first line. Leave small gaps so lines do not blur together.

Align cues to the picture. If the product appears at 2 seconds, start its line at 2.0. For a generated clip you did not time, watch it once and write the times down. A trim first can help: cut the clip to the length you want, then time the cues against the trimmed file, because the times are measured from the start of the video you send.

If the clip does have speech after all, skip the cues and send script_text so the captions follow the speaker. And if you work in Korean, the doc says Korean text on a Latin style is rejected, so pick a Hangul-capable style for those cues.

Cost and limits

Each accepted standalone caption job reserves and captures $0.20, which the doc says is for videos of up to 60 seconds under the current fixed estimate. Check the live price in the catalog. If you need captions on a longer Short, read captions on long video before you plan the batch.

Two ways the plan can go wrong. First, video_url has to be a public HTTPS URL. Second, burned-in captions are part of the picture, so viewers cannot turn them off. Keep a clean copy.

Caption job behaviour from Sume's Video captions doc (read 2026-10-06)
InputResult
Speech, no script_textSpeech-to-text, then burn
Speech plus script_textAlign your text to the speech
Silent clip, no cuesFails with caption_no_speech
Silent clip plus cuesBurns authored overlay text

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume