script_text, words, cues or segments: which caption input to send

Sume captions take only one of script_text, words, cues, segments. Your pick decides whether speech-to-text runs and what happens on a silent clip.

5 min readSume
All posts

Send exactly one of script_text, words, cues and segments to POST /v1/video-captions, or none. The Sume docs state you can send only one of them. The choice decides whether speech-to-text runs: with script_text it does, and Sume aligns your text to its word timings; with words, cues or segments it does not, and your text is burned at the times you give.

The four inputs

cues and segments are listed together in the docs as the phrase-level form.

Caption text inputs on Sume video captions (read 2026-10-05)
InputSpeech-to-text runsShapeUse it when
noneYesOnly video_urlThe transcript can be taken as spoken.
script_textYes, kept as the source of truth for timeOne stringYou have the exact script and the clip has speech.
wordsNoWord-level, each with text, start, end in secondsYou have word times, for example from a prior transcript, or you need to correct text.
cues or segmentsNoPhrase-level overlay cards with text, start, endSilent clips and authored overlay text.

Silent clips

Speech-to-captions works only when the clip has audible speech. On a silent clip the job fails with caption_no_speech, with next_action: use_overlay_captions. This is not a generic policy rejection. The fix is to send cues or segments with text, start and end, so Sume burns authored text without speech-to-text.

{
  "video_url": "https://example.com/silent.mp4",
  "cues": [
    { "text": "New this week", "start": 0.0, "end": 1.8 },
    { "text": "Free shipping", "start": 1.8, "end": 3.5 }
  ]
}

When script_text can fail

script_text keeps the speech-to-text word timings as the source of truth and aligns your text to them. The alignment can fail with typed public job errors, which the video captions page lists. If your script and the speech differ a lot, expect failure, not a silent fallback. A language hint helps speech-to-text, and language never selects the style or the font.

Restyle without redoing the transcript

To burn the same video with another style, send source_caption_id and not video_url. Sume reuses the stored word timings, and you pass words only to correct text.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume