Burn captions: pick one of script_text, words, cues or segments

Sume video captions accepts one wording source per job: script_text, words, cues or segments. When each fits and which ones skip speech-to-text.

4 min readSume
All posts

A Sume video caption job takes its wording from one of four places, and the docs, read 2026-10-05, say you can send only one of script_text, words, cues and segments. The choice decides whether speech-to-text runs, so it is the first thing to settle.

The price does not depend on it: each accepted standalone caption job reserves and captures $0.20 for a video up to 60 seconds.

Which one when

From the Sume video captions docs, read 2026-10-05
FieldTiming fromSTT runs?Fits
none of themSpeech-to-textYesSpeech clip, default wording
script_textSpeech-to-text, wording aligned to your scriptYesExact spelling of names and claims
wordsYou (word level)NoCorrected transcript you already have
cues or segmentsYou (phrase level, text, start, end)NoSilent clip or authored overlay text

Failure modes

  • script_text can fail alignment with script_alignment_mismatch or script_alignment_failed. The suggested next action is simplify_script_text_or_omit.
  • A speech-based job on a silent clip fails with caption_no_speech, and next_action points to use_overlay_captions.
  • Korean text on a Latin style such as slam returns 400 caption_hangul_text_latin_style.

Authored cues on a silent clip

{
  "video_url": "https://media.sume.com/artifacts/example/silent.mp4",
  "style": "punch",
  "cues": [
    {"text": "New: free shipping", "start": 0.5, "end": 2.5},
    {"text": "Ends Sunday", "start": 2.5, "end": 4.5}
  ]
}

To change only the look afterwards, send source_caption_id and a new style. Sume reuses the source video and stored word timings, so speech-to-text does not run twice.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume