Caption a silent product video: caption_no_speech and cues
A silent clip fails speech-to-captions with caption_no_speech. Send cues with text, start and end to burn authored overlay text on /v1/video-captions instead.

A silent product video cannot be captioned by speech recognition, so POST /v1/video-captions fails it with caption_no_speech and next_action: use_overlay_captions. The fix is to send cues (or segments), each with text, start and end. Sume then burns your authored text onto the clip without running speech-to-text.
Why the error is not a policy block
Speech-to-captions, whether it uses your script_text or finds the words itself, works only when the clip has audible speech. A product turntable, a sizzle shot or a music-only reel has none. The docs say this error is not a generic policy rejection. It tells you to pick another input, and it names the way out.
This is the usual failure when an ad pipeline runs a caption step after every generation. Check whether the clip has speech before you pick a mode.
Send authored cues
Overlay cues are your own lines on a timeline. The request needs the public HTTPS video_url and the cue list. Style, font and language remain optional.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: turntable-cues-001" \
-d '{
"video_url": "https://example.com/turntable.mp4",
"style": "slam",
"cues": [
{"text": "Cast iron, pre-seasoned", "start": 0.5, "end": 2.5},
{"text": "Oven safe", "start": 3.0, "end": 4.5}
]
}'Pick a style
If you omit style, the text sets it: slam for Latin text and black-outline for Korean. You can also pick punch, tiktok-green or korean-ad. A design object overrides one token at a time, such as the active colour, for a single request.
language only hints speech-to-text. It does not pick the style or the font, and with authored cues there is no speech-to-text to hint.
| Clip | Input to send | Result |
|---|---|---|
| Has speech | Nothing extra, or script_text | Speech-to-captions |
| Silent | cues or segments with text, start, end | Authored overlay text |
| Silent, no cues | Nothing | caption_no_speech |
Keep cues short
Silent social video is read, not heard. Keep a cue to a few words and about two seconds on screen. Use the first cue within the first second so the hook lands before a viewer scrolls.
Sources
Related posts
More in Media tools
- Remove the card behind burned-in captions: colors.card null
Set design.colors.card to null on a Sume caption job to show no card. It works on styles that support design, not on punch or tiktok-green.
- Captions with no style set: Korean gets black-outline, no gold tint
If you send no style, Sume picks one from the caption text, and the spoken word gets no gold tint. Set design.colors.active to add one.
- Cheapest product video ad: six stills in a Timeline render for $0.30
Six product photos become a 15 s vertical ad with no video model: a $0.10 Timeline render plus $0.20 caption cues. $0.30 against $3.43 for five Wan clips.
- Check an AI pattern tiles seamlessly with ImageChops.offset
Shift a Sume pattern image by half its size with ImageChops.offset and measure the seam stripe. A pass or fail score before you ship a repeating background.
Written by Sume