Add captions to a silent clip: cues instead of speech recognition

A clip with no speech fails caption_no_speech. Send cues with your own text and times to burn on-screen lines on silent or music-only video for $0.20.

5 min readSume
All posts

Micro-dramas, product loops and music-only Shorts often have no speech to transcribe. Reported by HeyOrca (read 2026-10-07), micro-dramas are rising across TikTok, Reels and Shorts, and many of them tell the story with on-screen lines. A caption job that runs speech recognition on a silent clip fails with caption_no_speech.

The fix is to give Video captions the text and the timing yourself, so no speech step runs.

Use cues

cues is a list of phrase cards with text, start and end in seconds. Send exactly one of script_text, words, cues and segments. The video is a public HTTPS video_url of up to 60 seconds, and the job is $0.20 (read 2026-10-07).

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: silent-caps-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/example/loop.mp4", "style": "punch", "cues": [{"text": "He never called back.", "start": 0.5, "end": 2.5}, {"text": "Until tonight.", "start": 3.0, "end": 4.5}]}'

Which input for which clip

Caption inputs for different clips, read 2026-10-07
ClipInput to sendWhy
Spoken, no scriptvideo_url onlySume transcribes and burns the words
Spoken, you have the scriptscript_textText is aligned to the speech timing
Spoken, words already timedwordsNo STT; your timings are used
Silent or music onlycues or segmentsNo speech to find; your times are used

Check the timing

Keep every cue inside the video's length and keep end after start; out-of-range values are rejected with a 400 rather than clamped. Leave a gap between cards, because overlapping cues compete on screen. Preview one frame at each cue before you caption a batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume