caption_no_speech on a silent clip: burn Halloween text with cues
A silent AI clip fails video captions with caption_no_speech. Pass cues with text, start and end instead: a worked giveaway announcement on Sume at $0.20.

If POST /v1/video-captions fails with caption_no_speech, your clip has no audible speech for the transcriber to hear. The fix is not a retry. Send cues (or its alias segments) with text, start and end in seconds, which burns exactly your copy at exactly your times and skips speech-to-text. The error itself says so: next_action is use_overlay_captions. The job costs $0.20 whether it is speech or overlay.
Why a Halloween clip trips it
Seasonal clips are often wordless: a fog shot, a pumpkin turning, a house with one lit window. A music-only bed or a model clip with ambient sound has nothing to transcribe. The standalone caption endpoint documents this case explicitly: a silent clip fails as caption_no_speech, not as a generic policy reject, so you can branch on it in code.
The request
Here is a 12-second giveaway announcement. Each cue is one on-screen phrase; the endpoint accepts up to 200 cues per job. Keep each card under about six words and give it at least a second on screen. words, cues, segments and script_text are mutually exclusive, so pick one.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: giveaway-captions-001" \
-d '{
"video_url": "https://example.com/fog-house.mp4",
"style": "slam",
"cues": [
{ "text": "TRICK OR TREAT", "start": 0.5, "end": 3.0 },
{ "text": "Win the pumpkin box", "start": 3.5, "end": 7.0 },
{ "text": "Comment to enter", "start": 7.5, "end": 11.5 }
]
}'Choices that matter
| Field | What it does |
|---|---|
| style | Omit for slam on Latin text; named looks include punch and tiktok-green |
| design | Per-request colours, placement and phrasing; not supported on punch or tiktok-green |
| cues | Phrase cards, text plus start and end seconds; skips speech-to-text |
| language | Speech hint only; it never selects the style or font |
Make cards read on a dark clip
Halloween footage is dark, which helps white text but can bury a thin outline. Use design.colors.active for a single accent such as orange, and raise typography.font_size_ratio rather than adding more words. If the source is bright, darken it first with the video filter dim op, covered in this post on dimming before captions.
Limits
Cues are timings you wrote, and Sume does not check that they match what happens on screen. The flat $0.20 covers videos up to 60 seconds under the current fixed estimate, so confirm live pricing in GET /v1/catalog before a long batch. The source must be a fetchable public HTTPS URL; private or signed URLs are rejected. Cues are not an SRT import, which is unsupported.
Sources
Related posts
More in Media tools
- Captions unreadable on busy footage: dim the clip, then burn
Busy B-roll can swallow burned captions. Sume can dim the clip with video-filter, then burn captions with colour overrides, for about $0.22 a clip.
- Check character drift across AI shots with video_frames
Pull up to 24 evenly spaced stills from each AI clip with the unbilled video_frames route and compare faces and products before you stitch shots together.
- Check an AI take says your script: hypit align and unmatched words
hypit align pairs each script token with the transcript words of a generated take, and lists unmatched words. What it measures and what it does not do.
- Join voiceover takes into one gapless track with Timeline audio
Concatenate up to 20 Sume-hosted voice takes with Timeline audio, with no seam silence and no re-synthesis, and re-base video starts from the returned offsets.
Written by Sume