Caption job failed: caption_no_speech on a music-only AI video

Sume's caption job fails with caption_no_speech when a clip has no audible speech. Send cues with text, start and end instead; it is the same $0.20.

4 min readSume
All posts

caption_no_speech means Sume's caption job tried speech-to-text on a clip with no audible speech, such as a music-only or silent AI video. Fix it by sending cues (or segments) with text, start, and end in seconds; Sume then burns your text without running speech-to-text, at the same $0.20 per job for clips up to 60 seconds.

The error is not a policy rejection. The docs say it carries next_action: use_overlay_captions, which is Sume telling you what to do next.

Why it happens

Video captions says speech-to-captions works only when the clip has audible speech, whether you pass script_text or let STT run. Many AI clips, especially wordless B-roll, have a music bed or ambient sound and no speech, so the job has nothing to align.

You can find this out before paying for the caption job. Video inspect with frames: false returns only the probe, and probe.has_audio tells you whether there is an audio track at all. A track with music still passes has_audio, so a listen or a transcript is the real test.

The cues request

The request takes the same video_url, a style, and a cues array. You can send only one of script_text, words, cues, and segments. Text lands exactly where you put it.

TikTok's own captions feature is a separate thing: the 2022 newsroom post describes auto-generated closed captions that viewers can turn on, and Sume's burned-in cues are separate from it.

Caption input forms on Sume (read 2026-10-07)
FieldNeeds speech?Use for
(none) / script_textYesTalking clips; STT sets the timing
wordsNoWord-level text you already have
cues / segmentsNoPhrase-level overlay text on silent clips
curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caption-cues-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/broll.mp4",
    "style": "punch",
    "cues": [
      {"text": "Three days earlier", "start": 0.0, "end": 1.8},
      {"text": "She had no idea", "start": 1.8, "end": 3.6}
    ]
  }'

Two details that save a retry

Cues must fit the style: slam, punch, and tiktok-green use Latin display faces, so Korean text on them returns 400 caption_hangul_text_latin_style. And punch and tiktok-green do not support design overrides.

Send an Idempotency-Key so a network retry does not bill twice, and poll GET /v1/jobs/{id}/status as described in Jobs and results.

Fix it in four steps

The error is not a bug in your request. It means the caption job found no speech to transcribe. Music-only clips, silent product shots, and clips with only sound effects all trigger it, and Sume's response points to use_overlay_captions as the next action.

Work through the fix in order. First, confirm with video-inspect that probe.has_audio is true or false and, if it is true, whether a transcript comes back. Second, if the clip really has no speech, write your own lines and send them as cues with start and end times. Third, if the clip has speech that the transcriber missed, send script_text so Sume aligns your text to the audio. Fourth, restyle later with source_caption_id instead of paying for a new transcription.

Remember the rule that cues, segments, words, and script_text are mutually exclusive: send only one per job. Each standalone caption job costs $0.20 for clips up to 60 seconds, so check the input first rather than retrying blindly.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume