Five caption looks on one clip: one transcription, four restyles
Test five caption styles on one video for $1.00: the first job transcribes, four source_caption_id restyles reuse the word timings. Request bodies included.

Testing five caption looks on one video costs $1.00 on Sume if every job is a standalone caption job at $0.20, and only the first job runs speech-to-text. The other four send source_caption_id instead of video_url, so Sume reuses the source video of that caption and the word timings that it already has. The docs state that the price does not change for a restyle, because a restyle is still a render.
The five looks
For English speech, the three Latin styles named in the docs are slam, punch and tiktok-green. slam also takes design overrides, so it gives you two more looks without a new style name. The docs say punch and tiktok-green do not support design.
Look 1 is the first job, with no style. For Latin text the docs say the default is slam. Looks 2 to 5 are restyles.
| Look | How you ask | Speech-to-text runs? | Price |
|---|---|---|---|
| 1. Default (slam for Latin text) | video_url, no style | yes | $0.20 |
| 2. punch | source_caption_id + style punch | no | $0.20 |
| 3. tiktok-green | source_caption_id + style tiktok-green | no | $0.20 |
| 4. slam, cyan emphasis | source_caption_id + style slam + design.colors.active | no | $0.20 |
| 5. slam, red emphasis | source_caption_id + style slam + design.colors.active | no | $0.20 |
| Total | 1 transcription | $1.00 |
The requests
The first request creates the caption and returns an id. The restyle bodies replace video_url with source_caption_id. Send a new Idempotency-Key for each restyle, so a retry of one look cannot be confused with another look.
# look 2: restyle the first caption as punch
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: look-punch-001" \
-d "{\"source_caption_id\": \"$CAPTION_ID\", \"style\": \"punch\"}"
# look 4: slam with a cyan emphasis color
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: look-slam-cyan-001" \
-d "{\"source_caption_id\": \"$CAPTION_ID\", \"style\": \"slam\", \"design\": {\"colors\": {\"active\": \"#22D3EE\"}}}"What a restyle does not fix
A restyle keeps the word timings, so it also keeps any transcription mistake. If the speech-to-text misspelled a product name, send words with the restyle to correct the text. Do not send words together with script_text, cues or segments: you can send only one of those four.
A restyle also cannot rescue a wrong style for the language. Korean text on slam, punch or tiktok-green returns 400 with caption_hangul_text_latin_style, so Korean clips need one of the Hangul identities, which the docs list in a separate table. Colors in design must be hex, rgb() or rgba(), or transparent. Other CSS syntax is rejected at request time, which means a typo in a color costs you nothing.
How to judge the five looks
Pick the criteria before you render, so that the test ends with a choice. For a phone feed, check three things on each output: whether the spoken word stands out from the rest of the phrase, whether a two-line phrase stays inside the frame, and whether the emphasis color still reads on your brightest scene. If the footage is too bright for any look, a video filter dim pass at $0.02 per job is a cheaper fix than a sixth caption job, but the dimmed clip is a new source and needs a new video_url job.
Keep the id of the winning caption. A later restyle on a different clip needs its own first job, because source_caption_id points at the source video and word timings of that one caption.
The bill
Five looks at $0.20 each is $1.00. If you had sent video_url five times instead, you would also pay $1.00 on the price list, but you would run speech-to-text five times and could get five slightly different transcripts. The restyle path keeps one transcript across all five looks, which is what you want when the test is about the look and not about the words.
Sources
Related posts
More in Media tools
- video-frames fps limit: 24 frames, fps up to 2, what it covers
Sume video-frames takes fps from just above 0 up to 2 and caps each call at 24 frames at mid-bin times. Table of sample times and the span 24 frames reach.
- Video filter /check is free: validate 8 ops before you pay $0.02
POST /v1/video-filter/check runs the same validation as the encode and bills nothing. See what it catches, the 8-op limit, and when a valid program still fails.
- Omni edit loop on Sume: three instructions, three jobs, about $3.00
Editing a clip with Gemini Omni one instruction at a time is one Sume job per pass. If each pass returns 8 s at 720p, three passes cost $3.00.
- Group TTS sentences into 5 to 14.8 second lip-sync segments in Python
A short Python planner that groups sentence durations into segments the Sume H3 Max lip-sync route accepts (5 to 14.8 s) and prints the 768p cost of each.
Written by Sume