Restyle burned-in captions without transcribing twice
Pass source_caption_id instead of video_url to re-burn a video under a new Sume caption style. Word timings are reused, so no second speech-to-text runs.

To change the look of captions you already made, send source_caption_id and a new style instead of a video_url. Sume reuses the earlier caption's source video and the word timings it already holds, so no second speech-to-text runs. The docs are clear that billing is unchanged: a restyle is still a render.
That makes it the right tool for A/B testing looks, and the wrong tool for saving money.
What does the request look like?
The body is a source_caption_id and the new style, for example { "source_caption_id": "...", "style": "black-outline" }. You get the id from the earlier caption job. Pass words alongside it only when you need to correct the wording.
| Restyle with source_caption_id | New job with video_url | |
|---|---|---|
| Speech-to-text | Not run again | Runs unless you pass cues, words or similar |
| Word timings | Reused | Recomputed |
| Wording fixes | Pass words | Pass script_text or cues |
| Billing | Unchanged, still a render | A render |
When should you use it?
Mind the script rule. Latin styles reject Korean text with a 400, so a restyle from a Hangul style to slam will fail for a Korean clip. Choose another Hangul style.
- Comparing two or three styles on the same clip before choosing a house look.
- Moving a Korean clip from one Hangul style to another, where the faces differ.
- Fixing a single misheard word without re-timing everything, by sending corrected
words. - Producing platform variants that differ only in caption design.
What about translated text?
A restyle keeps the original wording unless you correct it with words. For a translated version you would send new cues on the video instead, as in the SRT post, and check terms first with the glossary check.
A standalone caption job is priced at $0.20 for videos up to 60 seconds under the current estimate, and a restyle costs the same as a render does, so plan the number of variants. The full field list is in Video captions.
How do you keep track of variants?
Store the caption id, the style and any design overrides for every variant you render. When a client picks one, you can reproduce it, and when a style changes you know which jobs to redo. Remember that each restyle is billed as a render, so decide on two or three candidates first and not ten.
If you compare more than two styles, put the same cue on each at the same frame in a contact sheet. Differences in weight and placement show up faster side by side.
Sources
Related posts
More in Developers
- Resume an Omni batch after a crash with stable idempotency keys
A worker dies halfway through 40 Omni clips. Build Idempotency-Key from the item id so a rerun resubmits safely, and treat 409 and 429 as signals, not failures.
- Retrain a cloned voice without breaking old videos
HeyGen keeps a voice ID when a clone is retrained. On Sume, an avatar is referenced by a stable handle, and each text-to-speech job records its voice and model.
- Retry a failed image batch on Sume: not billed, same Idempotency-Key
On Sume a failed image generation is not billed, but a retry after a timeout can double-submit. Send an Idempotency-Key per image and retry only the submit.
- Retry hints in the body or a header: reading Sume's Retry-After
Notion repeats Retry-After in response bodies. Sume can send a retry-after header on 429s. Here is how to retry each Sume error code safely.
Written by Sume