Restyle burned-in captions without paying for a second transcription

Pass source_caption_id to POST /v1/video-captions to re-burn the same video in another style, reusing its word timings. Billing is still one render.

4 min readSume
All posts

To re-burn captions on a video in a different style, send source_caption_id and a new style to POST /v1/video-captions instead of a video_url. Sume reuses that caption's source video and its word timings, so no second speech-to-text runs. The docs are explicit that billing is unchanged: a restyle is still a render.

The request

The body is just the earlier caption id and a style. Styles available are slam, punch, tiktok-green, korean-ad, black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis. Add Idempotency-Key, as with other media jobs.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caption-restyle-clip-3-black-outline" \
  -d '{"source_caption_id": "CAPTION_ID", "style": "black-outline"}'

Facts

From the Video captions page, checked 2026-10-01.

Facts from docs.sume.com/models/video-captions, checked 2026-10-01
ItemValue
Restyle inputsource_caption_id instead of video_url
Second STTNone; word timings are reused
Correct wordingPass words alongside it
Standalone job price$0.20 for videos up to 60 s
Skip STT entirelySend words, cues or segments (mutually exclusive with script_text)
Typed errorscaption_no_speech, caption_hangul_text_latin_style

When to use it

Use it for A/B tests of caption looks on one clip, for brand variants of a single edit, and for a client who asks to see three styles. It saves the transcription step, which is where wording mistakes enter, so correct the wording once and pass the corrected words with the restyle request.

Silent clips have no speech to time. For those, supply cues or segments with text, start and end in seconds, and Sume burns exactly that copy at those times.

Keeping caption ids

Store the caption id with the clip when the first job finishes. Without it, a later restyle has to start from a video_url and transcribe again. A small table of clip, caption id, style and artifact URL is enough, and it doubles as the record of which look went with which export.

Limits

The restyle costs the same as a fresh caption job, so ten styles are ten renders. Two styles, weight-shift and korean-ad, animate the font weight axis that only the Pretendard face carries; on a static face they keep colour and scale emphasis but lose the weight travel. The first-run video_url must be a public HTTPS URL. Related: remove filler words.

Related posts

More in Media tools

All Media tools posts

Written by Sume