Caption a silent product video: caption_no_speech and cues

A silent clip fails speech-to-captions with caption_no_speech. Send cues with text, start and end to burn authored overlay text on /v1/video-captions instead.

5 min readSume
All posts

A silent product video cannot be captioned by speech recognition, so POST /v1/video-captions fails it with caption_no_speech and next_action: use_overlay_captions. The fix is to send cues (or segments), each with text, start and end. Sume then burns your authored text onto the clip without running speech-to-text.

Why the error is not a policy block

Speech-to-captions, whether it uses your script_text or finds the words itself, works only when the clip has audible speech. A product turntable, a sizzle shot or a music-only reel has none. The docs say this error is not a generic policy rejection. It tells you to pick another input, and it names the way out.

This is the usual failure when an ad pipeline runs a caption step after every generation. Check whether the clip has speech before you pick a mode.

Send authored cues

Overlay cues are your own lines on a timeline. The request needs the public HTTPS video_url and the cue list. Style, font and language remain optional.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: turntable-cues-001" \
  -d '{
    "video_url": "https://example.com/turntable.mp4",
    "style": "slam",
    "cues": [
      {"text": "Cast iron, pre-seasoned", "start": 0.5, "end": 2.5},
      {"text": "Oven safe", "start": 3.0, "end": 4.5}
    ]
  }'

Pick a style

If you omit style, the text sets it: slam for Latin text and black-outline for Korean. You can also pick punch, tiktok-green or korean-ad. A design object overrides one token at a time, such as the active colour, for a single request.

language only hints speech-to-text. It does not pick the style or the font, and with authored cues there is no speech-to-text to hint.

From the Sume video captions docs, read 2026-10-05
ClipInput to sendResult
Has speechNothing extra, or script_textSpeech-to-captions
Silentcues or segments with text, start, endAuthored overlay text
Silent, no cuesNothingcaption_no_speech

Keep cues short

Silent social video is read, not heard. Keep a cue to a few words and about two seconds on screen. Use the first cue within the first second so the hook lands before a viewer scrolls.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume