Restyle burned-in captions without paying for a second transcription
Pass source_caption_id to POST /v1/video-captions to re-burn the same video in another style, reusing its word timings. Billing is still one render.

To re-burn captions on a video in a different style, send source_caption_id and a new style to POST /v1/video-captions instead of a video_url. Sume reuses that caption's source video and its word timings, so no second speech-to-text runs. The docs are explicit that billing is unchanged: a restyle is still a render.
The request
The body is just the earlier caption id and a style. Styles available are slam, punch, tiktok-green, korean-ad, black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis. Add Idempotency-Key, as with other media jobs.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-restyle-clip-3-black-outline" \
-d '{"source_caption_id": "CAPTION_ID", "style": "black-outline"}'Facts
From the Video captions page, checked 2026-10-01.
| Item | Value |
|---|---|
| Restyle input | source_caption_id instead of video_url |
| Second STT | None; word timings are reused |
| Correct wording | Pass words alongside it |
| Standalone job price | $0.20 for videos up to 60 s |
| Skip STT entirely | Send words, cues or segments (mutually exclusive with script_text) |
| Typed errors | caption_no_speech, caption_hangul_text_latin_style |
When to use it
Use it for A/B tests of caption looks on one clip, for brand variants of a single edit, and for a client who asks to see three styles. It saves the transcription step, which is where wording mistakes enter, so correct the wording once and pass the corrected words with the restyle request.
Silent clips have no speech to time. For those, supply cues or segments with text, start and end in seconds, and Sume burns exactly that copy at those times.
Keeping caption ids
Store the caption id with the clip when the first job finishes. Without it, a later restyle has to start from a video_url and transcribe again. A small table of clip, caption id, style and artifact URL is enough, and it doubles as the record of which look went with which export.
Limits
The restyle costs the same as a fresh caption job, so ten styles are ten renders. Two styles, weight-shift and korean-ad, animate the font weight axis that only the Pretendard face carries; on a static face they keep colour and scale emphasis but lose the weight travel. The first-run video_url must be a public HTTPS URL. Related: remove filler words.
Related posts
More in Media tools
- Revert an AI video edit: why Sume trims never touch your source
Descript moved its Revert button next to the AI response. In an API pipeline revert is free: each Sume edit returns a new artifact and the source stays put.
- Runway Enhance Frame Rate: 10 target rates, 300 s, vs Sume
Runway Enhance Frame Rate takes clips up to 300 seconds at 1 credit per 2 seconds. Sume has no such model; its video-trim output.fps accepts 24, 25, 30 or 60.
- source_too_long_for_reference_ingest: clips over 300 s
Reference ingest rejects a clip over 300 seconds with source_too_long_for_reference_ingest. What the cap covers and the routes Sume's docs point to instead.
- Square YouTube Short (1:1) from a vertical master with Timeline
YouTube accepts square or vertical Shorts. Render a 1080x1080 version of a vertical clip with Sume Timeline, using fit blur or contain instead of a hard crop.
Written by Sume