Captions that match your voiceover script: script_text on Sume
Burned-in captions can follow your script instead of the speech-to-text guess. Sume's script_text aligns to STT timing; here are its failure codes and limits.

To make captions say exactly what your script says, pass script_text on POST /v1/video-captions. Sume keeps the speech-to-text word timings as the source of timing and aligns the burned-in wording to your script. So brand names, numbers and punctuation come from your text, and the timing comes from the audio.
Captions are on Google's mind too: Releasebot's Google page (read 2026-10-02) reports that the 2026-10-02 Workspace recap mentions customizable captions in Vids. Releasebot is a third-party page, so this is reported. Sume's own options are below.
What are my input choices?
script_text, words, cues and segments are mutually exclusive. Pick the one that matches what you have.
| You have | Pass | Speech-to-text runs? |
|---|---|---|
| A script and a spoken video | script_text | Yes, for timing |
| Word timings you trust | words | No |
| Phrase cards for a silent clip | cues or segments | No |
| Nothing | (omit) | Yes, its wording is burned |
What can go wrong?
Alignment can fail with a typed job error: script_alignment_mismatch or script_alignment_failed. The suggested next action is simplify_script_text_or_omit, so shorten the script or drop script_text and burn the STT wording. A silent clip fails as caption_no_speech; use cues or segments for that.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-script-001" \
-d '{"video_url": "https://media.sume.com/artifacts/example/clean.mp4", "style": "punch", "script_text": "Say hello to the Sume developer platform."}'Which style should I use?
Omit style and the wording decides: slam for Latin, black-outline for Korean. Latin styles such as slam, punch and tiktok-green have no Hangul glyphs, so Korean copy on them returns 400 with caption_hangul_text_latin_style. The design object overrides colors, typography, placement, phrasing and motion for one request, except on punch and tiktok-green.
What does it cost?
A standalone caption job is $0.20 for videos up to 60 seconds under the current fixed estimate; confirm it in GET /v1/catalog.
Why not just burn the STT wording?
Speech-to-text is a good first draft and a poor source of truth. It can mishear a product name, drop punctuation, or spell a number a different way than your legal team approved. When the script is the approved text, script_text makes the caption follow it while the audio still decides the timing.
If your script differs a lot from what was said, for example because the speaker ad-libbed, alignment can fail. Then the choice is to fix the script to match the recording, or to omit script_text and edit the STT output instead.
How do I restyle without paying for another transcription?
Pass source_caption_id with a new style. Sume reuses the earlier caption's source video and word timings, so no second speech-to-text runs. A restyle is still a render and is billed as one. Pass words alongside it only to correct the wording.
What should I check before I publish?
Watch the finished clip once with the sound off. Check that each caption matches the approved script, that no line is cut at the frame edge, and that the style reads on a phone. A caption job is a render, so catching a typo before the job costs less than re-running it afterwards.
Sources
Related posts
More in Media tools
- Reference audio for Wan 3.0 and MiniMax H3: WAV, MP3, 15 MB
Both Wan 3.0 and MiniMax H3 take WAV or MP3 reference audio up to 15 MB and 15 seconds in total. Where the clip counts differ and what Sume says.
- Wan 3.0 reference images: 20 MB, 240 to 8,000 px, 8:1 cap
Alibaba's Wan 3.0 reference images must be 20 MB or less, 240 to 8,000 px per side, ratio 8:1 at most. A check to run before wan-3.0 on Sume.
- Which AI video models take an end frame in Sume's video picker?
Wan 3.0, MiniMax H3, H3 Max and Auto take an end frame in Sume's Videos panel; Kling 3.0 and Grok Imagine do not. How to check the API list too.
- YouTube caption file for a Short: with timing or without timing?
YouTube's Upload file option asks for With timing or Without timing. Which to pick for a Short, and how Sume's transcript segments and burned-in captions fit.
Written by Sume