Caption inputs on Sume: script_text, words, cues or segments?

Four caption inputs and only one may be sent. When to use script_text with transcription, word timings, or phrase cues for a silent clip. Plus the errors.

5 min readSume
All posts

A Sume caption job accepts four ways to supply wording, and you may send only one: script_text, words, cues or segments. Use script_text when the clip has speech and you want your spelling; use words when you already have word timings; use cues or segments for phrase cards, including clips with no speech. Omit all four and Sume transcribes the audio and burns what it hears.

Which one for which job

The rules come from the Video captions guide and Sume's request validation, read 2026-10-03. script_text keeps speech-to-text as the timing source, so the clip needs audible speech. The other three skip transcription.

Caption inputs on POST /v1/video-captions, read 2026-10-03
InputNeeds speech in the clipTiming comes fromTypical use
noneYesSpeech-to-textQuick captions of a spoken clip
script_textYesSpeech-to-text, wording aligned to your scriptFix brand names
wordsNoYour word start and end timesTTS timings, or a restyle with corrections
cuesNoYour phrase start and end timesSilent clips, labels
segmentsNoSame as cuesSame as cues

What fails and why

Sending two inputs together returns a 400, because words, cues and segments all name the same timed wording and script_text has nothing to align onto if timings are given. A silent clip with no authored input fails as caption_no_speech. A script_text that cannot be matched to the speech fails as script_alignment_mismatch or script_alignment_failed.

Each phrase cue can hold up to 400 characters, and each word up to 200, so keep cards short enough to read.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume