Caption job failed: caption_no_speech on a music-only AI video
Sume's caption job fails with caption_no_speech when a clip has no audible speech. Send cues with text, start and end instead; it is the same $0.20.

caption_no_speech means Sume's caption job tried speech-to-text on a clip with no audible speech, such as a music-only or silent AI video. Fix it by sending cues (or segments) with text, start, and end in seconds; Sume then burns your text without running speech-to-text, at the same $0.20 per job for clips up to 60 seconds.
The error is not a policy rejection. The docs say it carries next_action: use_overlay_captions, which is Sume telling you what to do next.
Why it happens
Video captions says speech-to-captions works only when the clip has audible speech, whether you pass script_text or let STT run. Many AI clips, especially wordless B-roll, have a music bed or ambient sound and no speech, so the job has nothing to align.
You can find this out before paying for the caption job. Video inspect with frames: false returns only the probe, and probe.has_audio tells you whether there is an audio track at all. A track with music still passes has_audio, so a listen or a transcript is the real test.
The cues request
The request takes the same video_url, a style, and a cues array. You can send only one of script_text, words, cues, and segments. Text lands exactly where you put it.
TikTok's own captions feature is a separate thing: the 2022 newsroom post describes auto-generated closed captions that viewers can turn on, and Sume's burned-in cues are separate from it.
| Field | Needs speech? | Use for |
|---|---|---|
(none) / script_text | Yes | Talking clips; STT sets the timing |
words | No | Word-level text you already have |
cues / segments | No | Phrase-level overlay text on silent clips |
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-cues-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/broll.mp4",
"style": "punch",
"cues": [
{"text": "Three days earlier", "start": 0.0, "end": 1.8},
{"text": "She had no idea", "start": 1.8, "end": 3.6}
]
}'Two details that save a retry
Cues must fit the style: slam, punch, and tiktok-green use Latin display faces, so Korean text on them returns 400 caption_hangul_text_latin_style. And punch and tiktok-green do not support design overrides.
Send an Idempotency-Key so a network retry does not bill twice, and poll GET /v1/jobs/{id}/status as described in Jobs and results.
Fix it in four steps
The error is not a bug in your request. It means the caption job found no speech to transcribe. Music-only clips, silent product shots, and clips with only sound effects all trigger it, and Sume's response points to use_overlay_captions as the next action.
Work through the fix in order. First, confirm with video-inspect that probe.has_audio is true or false and, if it is true, whether a transcript comes back. Second, if the clip really has no speech, write your own lines and send them as cues with start and end times. Third, if the clip has speech that the transcriber missed, send script_text so Sume aligns your text to the audio. Fourth, restyle later with source_caption_id instead of paying for a new transcription.
Remember the rule that cues, segments, words, and script_text are mutually exclusive: send only one per job. Each standalone caption job costs $0.20 for clips up to 60 seconds, so check the input first rather than retrying blindly.
Sources
Related posts
More in Developers
- Captions out of sync with the audio: check STT word times and offsets
Captions running early or late usually trace to an unapplied offset. How Sume STT word times work, which offset to add, and a Python merge that applies it.
- Chain TikTok mini-drama episodes: last frame in, first frame out
Pull the final still of episode one with Sume video frames, then send it as first_frame for episode two so the next shot starts where the last ended.
- Chapter timestamps for narrated audio from concat segment offsets
Join one TTS file per chapter with timeline audio concat, then turn the returned segments[] start offsets into mm:ss chapter lines with a short Python script.
- Check a reference image URL before sending it to Sume Images
Sume rejects localhost, private-network and non-HTTPS reference URLs before submission. A short Python pre-flight check that catches them first.
Written by Sume