caption_no_speech: why a caption job fails on a silent clip

A Sume caption job returns caption_no_speech when the clip has no audible speech. The fix is cues, segments or words, not a retry.

5 min readSume
All posts

caption_no_speech means the clip you sent has no audible speech, so speech-to-text had nothing to turn into captions. Sume's job result carries next_action: use_overlay_captions. Retrying the same request will fail the same way. To burn text onto a silent clip, send your own cues (or segments) with text, start and end, and Sume skips speech-to-text.

What triggers it and what it is not

Speech-to-captions runs when you send script_text, or when you send nothing and let speech-to-text decide the words. Both paths need audio with speech. A music-only reel, a screen recording with the mic off, or a clip whose audio track is empty fails the same way.

The docs state this is not a generic policy rejection. Your video was allowed; it simply has nothing to transcribe. That matters for how you handle it: do not rewrite the prompt or change the style, change the input form.

Four input forms and when each applies

You may send only one of script_text, words, cues and segments on a request.

Caption input forms, read 2026-10-08 from the Sume docs
FieldNeeds speech?Use it when
none (speech-to-text)YesThe clip has a voice and you accept the transcript
script_textYesYou have the script and want it aligned to the spoken words
wordsNoYou have word-level times (text, start, end in seconds)
cues / segmentsNoYou have phrase-level overlay text for a silent clip

Check before you pay

Probe the clip first. A video_inspect call with frames: false returns probe.has_audio, and the inspect docs say a silent clip with transcribe: true fails with inspect_source_has_no_audio. One probe is cheaper than a failed caption job, and it tells you which input form to build.

If has_audio is true but the job still returns caption_no_speech, the audio is probably music or room tone. Send cues for that clip, or detach the audio and listen to it.

A cues request

The cue times are in seconds on the video timeline. Sume burns the text exactly as written at those times.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: overlay-cues-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/silent.mp4",
    "cues": [
      {"text": "Order by Friday", "start": 0.5, "end": 2.5},
      {"text": "Ships in 2 days", "start": 2.5, "end": 5.0}
    ]
  }'

Where it fits in a pipeline

Treat caption_no_speech as a routing signal. A pipeline that captions mixed footage can run a free-ish probe on each clip, then send speech clips down the speech path and silent clips down the overlay path. That is cleaner than catching the error afterward, although catching it works too since next_action tells your code what to do.

Overlay text is yours to write. Sume burns exactly the text and times you send, so a silent product loop can carry a headline in the first two seconds and a call to action in the last three. Choose a Hangul style for Korean overlay text, because the Latin styles reject Hangul with a 400.

  • Silent clip plus cues: no speech-to-text, text burned at your times.
  • Clip with speech: do nothing special, send video_url alone.
  • Unsure: probe with frames: false and read probe.has_audio.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume