Captions that match your voiceover script: script_text on Sume

Burned-in captions can follow your script instead of the speech-to-text guess. Sume's script_text aligns to STT timing; here are its failure codes and limits.

4 min readSume
All posts

To make captions say exactly what your script says, pass script_text on POST /v1/video-captions. Sume keeps the speech-to-text word timings as the source of timing and aligns the burned-in wording to your script. So brand names, numbers and punctuation come from your text, and the timing comes from the audio.

Captions are on Google's mind too: Releasebot's Google page (read 2026-10-02) reports that the 2026-10-02 Workspace recap mentions customizable captions in Vids. Releasebot is a third-party page, so this is reported. Sume's own options are below.

What are my input choices?

script_text, words, cues and segments are mutually exclusive. Pick the one that matches what you have.

Caption inputs. Sume Video captions docs, read 2026-10-02.
You havePassSpeech-to-text runs?
A script and a spoken videoscript_textYes, for timing
Word timings you trustwordsNo
Phrase cards for a silent clipcues or segmentsNo
Nothing(omit)Yes, its wording is burned

What can go wrong?

Alignment can fail with a typed job error: script_alignment_mismatch or script_alignment_failed. The suggested next action is simplify_script_text_or_omit, so shorten the script or drop script_text and burn the STT wording. A silent clip fails as caption_no_speech; use cues or segments for that.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caption-script-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/example/clean.mp4", "style": "punch", "script_text": "Say hello to the Sume developer platform."}'

Which style should I use?

Omit style and the wording decides: slam for Latin, black-outline for Korean. Latin styles such as slam, punch and tiktok-green have no Hangul glyphs, so Korean copy on them returns 400 with caption_hangul_text_latin_style. The design object overrides colors, typography, placement, phrasing and motion for one request, except on punch and tiktok-green.

What does it cost?

A standalone caption job is $0.20 for videos up to 60 seconds under the current fixed estimate; confirm it in GET /v1/catalog.

Why not just burn the STT wording?

Speech-to-text is a good first draft and a poor source of truth. It can mishear a product name, drop punctuation, or spell a number a different way than your legal team approved. When the script is the approved text, script_text makes the caption follow it while the audio still decides the timing.

If your script differs a lot from what was said, for example because the speaker ad-libbed, alignment can fail. Then the choice is to fix the script to match the recording, or to omit script_text and edit the STT output instead.

How do I restyle without paying for another transcription?

Pass source_caption_id with a new style. Sume reuses the earlier caption's source video and word timings, so no second speech-to-text runs. A restyle is still a render and is billed as one. Pass words alongside it only to correct the wording.

What should I check before I publish?

Watch the finished clip once with the sound off. Check that each caption matches the approved script, that no line is cut at the frame edge, and that the style reads on a phone. A caption job is a render, so catching a typo before the job costs less than re-running it afterwards.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume