caption_no_speech: caption a silent clip with cues

caption_no_speech means the clip has no audible speech. Sume's fix is next_action use_overlay_captions: send cues with text, start and end and skip ASR.

4 min readSume
All posts

caption_no_speech is the failure Sume gives when a caption job has no audible speech to transcribe. It carries next_action: use_overlay_captions: resend the job with cues (or segments), each with text, start and end, and Sume burns that copy without running speech-to-text.

Auto captions of the kind Descript redesigned on 2026-09-17 assume someone is talking. The rule here is in Video captions.

What does the cues body look like?

Times are in seconds. This body burns two overlay cards on a silent clip.

{
  "video_url": "https://media.sume.com/artifacts/example/silent.mp4",
  "cues": [
    { "text": "Spring collection", "start": 0.0, "end": 2.0 },
    { "text": "Out now", "start": 2.0, "end": 4.0 }
  ]
}

Which inputs can I combine?

Only one wording source per request.

Caption wording inputs, from the Sume docs read 2026-09-30.
InputRuns speech-to-text?Use it for
noneYesBurn what is spoken
script_textYes, for timingsAlign your wording to the speech
wordsNoWord-level timing you already have
cues / segmentsNoPhrase-level overlay cards, silent clips

Can I upload an SRT instead?

No. The docs say SRT uploads and provider task ids are unsupported; pass phrase-level text as cues or segments. script_text, words, cues and segments are mutually exclusive, so a request with two of them is invalid.

When does this matter for ads?

A clip with music and no speech still needs words on screen for viewers who watch with the sound off. See captions on Facebook feed video ads for that context, then author the cues yourself.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume