Captions fail on a silent Short with caption_no_speech: send cues
A silent clip has no speech to transcribe, so Sume's captions API returns caption_no_speech. Send cues with text, start and end to burn overlay text instead.

A silent clip fails the Sume captions API with caption_no_speech, because speech-to-text has nothing to read. The fix named in the error is use_overlay_captions: send cues with text, start and end, and Sume burns your text at those times without transcribing.
Why a silent clip fails
The behavior is in the video captions docs. It matters for Shorts because a lot of AI clips are silent or music-only, and a Short is often watched with the sound off.
What the error means
A job with no script_text, words, cues or segments runs speech-to-text. If the clip has no audible speech, the job fails with caption_no_speech and next_action: use_overlay_captions. The docs say this is not a generic policy rejection. You can send only one of script_text, words, cues and segments in a request.
| You send | Speech-to-text runs | Use it for |
|---|---|---|
| Nothing extra | Yes | Clips with audible speech |
script_text | Yes, for timing; text aligned to your script | Speech clips where you have the script |
words | No | Word-level text you timed yourself |
cues or segments | No | Phrase-level overlay on silent clips |
Probe before you call
Check first whether the clip has audio. A probe from video inspect with frames: false returns probe.has_audio, and a transcribe request on a clip with no audio gives inspect_source_has_no_audio.
Send cues
Time each cue in seconds. A 12-second silent Short with three lines could look like this:
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: silent-short-cues-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/silent.mp4",
"style": "slam",
"cues": [
{ "text": "No sound needed", "start": 0.0, "end": 3.0 },
{ "text": "Three steps", "start": 3.5, "end": 7.0 },
{ "text": "Try it today", "start": 8.0, "end": 12.0 }
]
}'Price and constraints
Each accepted standalone caption job reserves and captures $0.20, for clips of at most 60 seconds under the current fixed estimate. Confirm the live price at GET /v1/catalog. The video must be a public HTTPS URL that Sume can fetch.
Pick a style that matches the text
Latin text works on slam, punch and tiktok-green. Hangul text on those Latin styles is rejected with caption_hangul_text_latin_style, so match the style to the text.
Write cues that are readable
Cues are phrase cards, so keep each one short enough to read at a glance. A reasonable rule is a few words per card and a pause between cards, so the viewer sees one idea at a time. Because you author the cues, you control the phrasing directly.
Match cue times to the cuts
Do not leave gaps larger than you intend. If a cue ends at 3.0 and the next starts at 3.5, there is half a second with no text on screen, which can look like a glitch if the clip has a hard cut there. Align the cue boundaries with the visual cuts, and read the timestamps from stills if you are unsure. Video inspect with frames set to at returns stills at explicit times so you can see what is on screen at each cue.
Restyle without starting over
If you later want to try a different look on the same clip, you do not need to run the job from the original file. The docs describe restyling with source_caption_id, which reuses the source video and its timings instead of transcribing again. The price does not change because a restyle is still a render.
Fail early and cheap
A last practical point is the failure cost. Because the API refuses a bad request before it renders, a wrong style or a malformed cue gives a 400 and you do not pay for the render, according to the docs. Test a small clip first, then the full file, and keep the Idempotency-Key stable on retries so a repeated call returns the same job instead of creating a second one.
One branch in your script
Combine this with a probe step in your pipeline. If probe.has_audio is false, send cues. If it is true, let speech-to-text run. That one branch removes the most common cause of caption failures on AI clips, and it keeps the behavior predictable for whoever maintains the script later.
Input limits
Remember that the source must be reachable. The captions docs require a public HTTPS video URL that Sume can fetch, and they reject localhost, private-network, signed or private URLs and provider task URLs. SRT uploads are not supported, so convert subtitle files into cues yourself.
Where cues fit in a pipeline
Place the cue step after trimming and before the final render. Trim first, so the times you author match the final cut. If you trim after burning, every cue shifts and the text no longer lines up with the picture, and you pay for the caption job again.
The takeaway
Probe for audio, then send cues for silent clips and let speech-to-text handle the rest.
Sources
Related posts
More in Developers
- 16 AI shots in one Sume Timeline render: the 8-fade cap
A 16-shot cut fits one Timeline render, but fades are capped at 8 in a row and renders chunk past 12 slots. Plan the cuts, with the doc limits.
- Size a batch from generation_limits so no clip hits queue_full
Read accepted_generation_jobs_limit from a Sume submit response and slice your clips. A 50-clip batch leaves 2 for a later wave on Startup and 26 on Pro.
- Voice model updated in place with no API change: how to detect it
Nova 2 Sonic was refreshed in place in May with no API change. If a vendor can change your voice silently, log the model id and a canary clip. Sume code inside.
- spend_approval_queue_full 429: clear pending approvals, do not retry
A thread with too many pending spend approvals gets 429 spend_approval_queue_full. Resolve the pending ones first; a 503 store_misconfigured is for support.
Written by Sume