Five caption looks on one clip: one transcription, four restyles

Test five caption styles on one video for $1.00: the first job transcribes, four source_caption_id restyles reuse the word timings. Request bodies included.

5 min readSume
All posts

Testing five caption looks on one video costs $1.00 on Sume if every job is a standalone caption job at $0.20, and only the first job runs speech-to-text. The other four send source_caption_id instead of video_url, so Sume reuses the source video of that caption and the word timings that it already has. The docs state that the price does not change for a restyle, because a restyle is still a render.

The five looks

For English speech, the three Latin styles named in the docs are slam, punch and tiktok-green. slam also takes design overrides, so it gives you two more looks without a new style name. The docs say punch and tiktok-green do not support design.

Look 1 is the first job, with no style. For Latin text the docs say the default is slam. Looks 2 to 5 are restyles.

Plan for five looks on one clip (read 2026-10-08)
LookHow you askSpeech-to-text runs?Price
1. Default (slam for Latin text)video_url, no styleyes$0.20
2. punchsource_caption_id + style punchno$0.20
3. tiktok-greensource_caption_id + style tiktok-greenno$0.20
4. slam, cyan emphasissource_caption_id + style slam + design.colors.activeno$0.20
5. slam, red emphasissource_caption_id + style slam + design.colors.activeno$0.20
Total1 transcription$1.00

The requests

The first request creates the caption and returns an id. The restyle bodies replace video_url with source_caption_id. Send a new Idempotency-Key for each restyle, so a retry of one look cannot be confused with another look.

# look 2: restyle the first caption as punch
curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: look-punch-001" \
  -d "{\"source_caption_id\": \"$CAPTION_ID\", \"style\": \"punch\"}"

# look 4: slam with a cyan emphasis color
curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: look-slam-cyan-001" \
  -d "{\"source_caption_id\": \"$CAPTION_ID\", \"style\": \"slam\", \"design\": {\"colors\": {\"active\": \"#22D3EE\"}}}"

What a restyle does not fix

A restyle keeps the word timings, so it also keeps any transcription mistake. If the speech-to-text misspelled a product name, send words with the restyle to correct the text. Do not send words together with script_text, cues or segments: you can send only one of those four.

A restyle also cannot rescue a wrong style for the language. Korean text on slam, punch or tiktok-green returns 400 with caption_hangul_text_latin_style, so Korean clips need one of the Hangul identities, which the docs list in a separate table. Colors in design must be hex, rgb() or rgba(), or transparent. Other CSS syntax is rejected at request time, which means a typo in a color costs you nothing.

How to judge the five looks

Pick the criteria before you render, so that the test ends with a choice. For a phone feed, check three things on each output: whether the spoken word stands out from the rest of the phrase, whether a two-line phrase stays inside the frame, and whether the emphasis color still reads on your brightest scene. If the footage is too bright for any look, a video filter dim pass at $0.02 per job is a cheaper fix than a sixth caption job, but the dimmed clip is a new source and needs a new video_url job.

Keep the id of the winning caption. A later restyle on a different clip needs its own first job, because source_caption_id points at the source video and word timings of that one caption.

The bill

Five looks at $0.20 each is $1.00. If you had sent video_url five times instead, you would also pay $1.00 on the price list, but you would run speech-to-text five times and could get five slightly different transcripts. The restyle path keeps one transcript across all five looks, which is what you want when the test is about the look and not about the words.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume