Caption a silent AI video: fixing caption_no_speech
A silent clip fails POST /v1/video-captions with caption_no_speech. Send cues with text, start and end to burn authored captions without speech-to-text.

A video with no audible speech fails POST /v1/video-captions with caption_no_speech, because the default path runs speech-to-text. Pass cues (or segments) with text, start and end in seconds instead. Sume then burns exactly that copy at those times and skips speech-to-text.
Why a silent clip fails
Standalone captions take a public HTTPS video_url. With script_text, or with nothing, Sume needs audible speech to time the words. A silent clip returns a typed failure caption_no_speech with next_action: use_overlay_captions. It is not a generic policy rejection, and the fix is in the request, not the video.
Send cues instead
cues and segments are phrase-level overlay cards. words is word-level. script_text, words, cues and segments are mutually exclusive, so send one.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: silent-caption-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/silent.mp4",
"style": "slam",
"cues": [
{"text": "New drop", "start": 0.2, "end": 1.8},
{"text": "Out this Friday", "start": 2.0, "end": 4.0}
]
}'Style and language choices
Omit style and the wording decides: slam for Latin text, black-outline for Korean. You can also name slam, punch, tiktok-green or korean-ad. Korean copy on slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style, because those faces have no Hangul glyphs; pick a Hangul style for Korean cues.
design overrides colors, typography, placement, phrasing and motion for one request, but it is not supported on punch or tiktok-green.
Cost, limits and polling
Each accepted standalone caption job reserves and captures $0.20 for videos up to 60 seconds under the current fixed estimate; confirm live pricing in GET /v1/catalog. The video_url must be a fetchable public HTTPS URL. Localhost, private-network, signed or private URLs are rejected, and SRT uploads are not supported; send phrase text as cues instead.
Poll GET /v1/jobs/{id}/status and read GET /v1/video-captions/{id} when ready. To restyle the same video later, pass source_caption_id instead of video_url; billing is unchanged because a restyle is still a render.
| Input | Needs speech | Use when |
|---|---|---|
none or script_text | Yes | The clip has a voice |
words | No | You have word timings |
cues or segments | No | The clip is silent or you author the copy |
Sources
Related posts
More in Media tools
- Change the caption highlight colour without changing the style
Cheap transcription made captions routine; brand colour is what is left. One design field changes the spoken-word colour and keeps everything else in the style.
- Turn a music composition plan into a Sume time-range prompt (Python)
ElevenLabs music_v2_5 plans allow 6,132 characters in up to 30 lines. Sume's Music prompt takes 5000 characters. A Python converter for time ranges.
- Fix one wrong word in burned-in captions without a re-transcribe
One misheard word in a burned-in caption? Re-burn with source_caption_id and a words list instead of running speech-to-text again. Request body included.
- Cover or contain: fit a 16:9 clip into a vertical stack
Sume timeline compose takes video_fit cover, contain or stretch. Work out what each does to a 16:9 clip in the 1080x1920 default stack, then check one frame.
Written by Sume