Restyle captions four ways for 80 cents without transcribing twice
Each Sume caption render is $0.20, so four styles of one clip cost $0.80. A restyle with source_caption_id reuses the word timings and skips re-transcribing.

The cost and the shortcut
A restyle is still a render, so each of four style variants bills $0.20, for $0.80 in total. The saving is not on price but on correctness and time: send source_caption_id instead of video_url and Sume reuses the source video and the word timings it already has, so speech-to-text does not run a second time.
How to send it
The first job takes video_url. Later jobs take only source_caption_id and a style. Send words with it only to correct the text. You can send only one of script_text, words, cues and segments per request.
Available named styles include slam, punch, tiktok-green and korean-ad, plus black-outline as the Korean default, per the docs.
| Job | Body | Cost |
|---|---|---|
| 1 (original) | video_url, style slam | $0.20 |
| 2 | source_caption_id, style punch | $0.20 |
| 3 | source_caption_id, style tiktok-green | $0.20 |
| 4 | source_caption_id, style black-outline | $0.20 |
| Total | $0.80 |
Why this matters for correctness
If the first transcription misheard a product name, fix it once with words on the restyle, and the corrected text carries through the variants you create from that point. Running speech-to-text again on each variant could give different splits and different mistakes on each.
For a 20-clip campaign with 4 variants each, that is 80 renders, 80 x $0.20 = $16.00.
When not to restyle
Restyling cannot add speech to a silent clip; use cues for that. A style that moves emphasis by word weight looks different on a static face, as the docs note. Reference: Video captions.
Choosing which variants to make
Four variants is a test design, not a requirement. For social posts, one bold style for short clips and one clean outline for longer talking-head clips cover most cases. Render the pair on one clip, compare retention in your own analytics and then settle on a default for the series.
Every variant is a full render, so the cost grows linearly: six variants are $1.20.
Keeping the source
The restyle needs the first caption job's id, so store it. The docs describe the restyle as reusing that caption's source video and word timings, so do not delete the source asset. If the source is gone, plan to start again from a video_url at $0.20 and a fresh transcription.
Sources
More in Media tools
- Restyle captions three times with source_caption_id: $0.60 total
To test three caption looks on one video, send source_caption_id: no second transcription, but each restyle is a render at the same job price.
- script_alignment_mismatch: fixing a caption job with script_text
When script_text does not match the speech, a Sume caption job fails with script_alignment_mismatch or script_alignment_failed. Simplify the script or omit it.
- Six 5-second reference clips: $2.28 on H3 to $17.34 on Seedance 2.5
Cost of six 5-second 720p reference-to-video clips on Sume by model: MiniMax H3 $2.28 at 768p, Wan 3.0 $3.78, Seedance 2.0 Fast $9.12, Seedance 2.5 $17.34.
- Timeline soundtrack duck_db: 0 to 20 dB, needs a real audio spine
Sume Timeline soundtrack.duck_db accepts 0 to 20 and returns duck_requires_audio_spine on silence. Use it with a voiceover spine; gain_db covers a static level.
Written by Sume