Captions fail on a silent Short with caption_no_speech: send cues

A silent clip has no speech to transcribe, so Sume's captions API returns caption_no_speech. Send cues with text, start and end to burn overlay text instead.

5 min readSume
All posts

A silent clip fails the Sume captions API with caption_no_speech, because speech-to-text has nothing to read. The fix named in the error is use_overlay_captions: send cues with text, start and end, and Sume burns your text at those times without transcribing.

Why a silent clip fails

The behavior is in the video captions docs. It matters for Shorts because a lot of AI clips are silent or music-only, and a Short is often watched with the sound off.

What the error means

A job with no script_text, words, cues or segments runs speech-to-text. If the clip has no audible speech, the job fails with caption_no_speech and next_action: use_overlay_captions. The docs say this is not a generic policy rejection. You can send only one of script_text, words, cues and segments in a request.

Caption input forms, Sume video captions docs (read 2026-10-05)
You sendSpeech-to-text runsUse it for
Nothing extraYesClips with audible speech
script_textYes, for timing; text aligned to your scriptSpeech clips where you have the script
wordsNoWord-level text you timed yourself
cues or segmentsNoPhrase-level overlay on silent clips

Probe before you call

Check first whether the clip has audio. A probe from video inspect with frames: false returns probe.has_audio, and a transcribe request on a clip with no audio gives inspect_source_has_no_audio.

Send cues

Time each cue in seconds. A 12-second silent Short with three lines could look like this:

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: silent-short-cues-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/silent.mp4",
    "style": "slam",
    "cues": [
      { "text": "No sound needed", "start": 0.0, "end": 3.0 },
      { "text": "Three steps", "start": 3.5, "end": 7.0 },
      { "text": "Try it today", "start": 8.0, "end": 12.0 }
    ]
  }'

Price and constraints

Each accepted standalone caption job reserves and captures $0.20, for clips of at most 60 seconds under the current fixed estimate. Confirm the live price at GET /v1/catalog. The video must be a public HTTPS URL that Sume can fetch.

Pick a style that matches the text

Latin text works on slam, punch and tiktok-green. Hangul text on those Latin styles is rejected with caption_hangul_text_latin_style, so match the style to the text.

Write cues that are readable

Cues are phrase cards, so keep each one short enough to read at a glance. A reasonable rule is a few words per card and a pause between cards, so the viewer sees one idea at a time. Because you author the cues, you control the phrasing directly.

Match cue times to the cuts

Do not leave gaps larger than you intend. If a cue ends at 3.0 and the next starts at 3.5, there is half a second with no text on screen, which can look like a glitch if the clip has a hard cut there. Align the cue boundaries with the visual cuts, and read the timestamps from stills if you are unsure. Video inspect with frames set to at returns stills at explicit times so you can see what is on screen at each cue.

Restyle without starting over

If you later want to try a different look on the same clip, you do not need to run the job from the original file. The docs describe restyling with source_caption_id, which reuses the source video and its timings instead of transcribing again. The price does not change because a restyle is still a render.

Fail early and cheap

A last practical point is the failure cost. Because the API refuses a bad request before it renders, a wrong style or a malformed cue gives a 400 and you do not pay for the render, according to the docs. Test a small clip first, then the full file, and keep the Idempotency-Key stable on retries so a repeated call returns the same job instead of creating a second one.

One branch in your script

Combine this with a probe step in your pipeline. If probe.has_audio is false, send cues. If it is true, let speech-to-text run. That one branch removes the most common cause of caption failures on AI clips, and it keeps the behavior predictable for whoever maintains the script later.

Input limits

Remember that the source must be reachable. The captions docs require a public HTTPS video URL that Sume can fetch, and they reject localhost, private-network, signed or private URLs and provider task URLs. SRT uploads are not supported, so convert subtitle files into cues yourself.

Where cues fit in a pipeline

Place the cue step after trimming and before the final render. Trim first, so the times you author match the final cut. If you trim after burning, every cue shifts and the text no longer lines up with the picture, and you pay for the caption job again.

The takeaway

Probe for audio, then send cues for silent clips and let speech-to-text handle the rest.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume