Test several caption looks on one video without re-transcribing

Send source_caption_id with a new style to re-burn a video's captions in a different look. No second speech-to-text; each variant is still a $0.20 render.

4 min readSume
All posts

To test several caption looks on the same video, create the first caption job with video_url, then send source_caption_id with a different style for each variant. Sume reuses the source video and the word timings it already has, so speech-to-text does not run again. The price does not drop, because a restyle is still a render at $0.20 per job.

The styles to compare

Caption styles named in the docs (read 2026-10-07)
StyleNote from the docs
slamDefault for Latin text when no style is sent
punchSelectable by name
tiktok-greenSelectable by name
black-outlineDefault for Korean text when no style is sent
clip-wipeEach word wipes in left to right; the docs call it the clearest at small phone sizes
korean-adCapCut-style karaoke for Korean speech; use with language: "ko"

The request

The first job takes video_url and optionally style, font, language, design and one of script_text, words, cues or segments. Every later variant is { "source_caption_id": "...", "style": "black-outline" }. Send words with it only to correct the text.

If you do not send a style at all, the caption text picks one: slam for Latin text. If you pick a style yourself, Sume renders that style, and a Latin style with Hangul text is a 400, not a silent fallback.

Fine tune with design overrides

design overrides the style's tokens for one request: colors, typography, placement, phrasing and motion. One key changes one thing, so an accent color test is a one-key request rather than a new style. If you do not send a style, the render has no accent color at all; a default style keeps its own gold tint.

A practical test plan is to change one thing at a time. Keep the text and timing fixed, vary only the style across the first round, then keep the winner and vary one design token in the second round. Because every variant is its own job with its own Idempotency-Key, you can run them in parallel and discard the ones you do not publish.

What a test costs

Four looks on a clip of 60 seconds or less is one first job plus three restyles, four renders at $0.20, so $0.80. Give each its own Idempotency-Key, poll each job, and compare the results on the platform you publish to. Sume tells you what rendered, not which look performs, so the measurement is yours.

If you already have a transcript from another source, skip speech-to-text entirely: send words for word-level timing or cues and segments for phrase-level cards. That is also the route for silent clips, where there is nothing to transcribe. Only one of script_text, words, cues and segments is accepted in a request.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume