Change the caption highlight colour without changing the style
Cheap transcription made captions routine; brand colour is what is left. One design field changes the spoken-word colour and keeps everything else in the style.

How do you change the highlighted word colour in burned-in captions without redesigning them? Send a design object with colors.active. It merges over the style you named, so one key changes one thing and everything else stays as the style defined it.
Transcription keeps getting cheaper: Microsoft's streaming model (read 2026-10-04) is priced at $0.54 per hour of audio through year-end. When timing is a commodity, the part viewers notice is the look, and for a brand that starts with the accent colour.
The request
Required for a standalone caption job is video_url, a public HTTPS URL. This example burns the black-outline style with a cyan spoken word, as in the video captions guide.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-cyan-001" \
-d '{
"video_url": "https://example.com/clean.mp4",
"style": "black-outline",
"design": { "colors": { "active": "#22D3EE" } }
}'
What you can and cannot set
A style is a set of design tokens and design overrides them for one request. Every field is optional.
| Group | Fields |
|---|---|
| colors | base, active, stroke, accent, accent_deep, card (null draws no card) |
| typography | base_weight, active_weight, active_scale, font_size_ratio, safe_width_ratio, stroke_width_px |
| placement | anchor_ratio, landscape_anchor_ratio |
| phrasing | max_words, max_chars, pause_seconds |
| motion | enter_seconds, exit_seconds, emphasis_in_seconds, emphasis_out_seconds |
Rules that save a retry
Colours are hex, rgb()/rgba() or transparent; other CSS syntax is rejected rather than drawn into the render. Numbers outside their documented range return 400, so a bad look fails at request time instead of rendering wrong and billing.
design is not supported on punch or tiktok-green; those render on a path that reads none of these tokens.
An omitted style drops the gold spoken-word tint both defaults were authored with, so a defaulted render keeps the fill colour on the spoken word. Naming black-outline or slam keeps that style's gold tint, and design.colors.active sets the colour either way.
A standalone caption job reserves $0.20 for videos up to 60 seconds under the current fixed estimate; confirm live pricing in GET /v1/catalog.
Sources
Related posts
More in Media tools
- Turn a music composition plan into a Sume time-range prompt (Python)
ElevenLabs music_v2_5 plans allow 6,132 characters in up to 30 lines. Sume's Music prompt takes 5000 characters. A Python converter for time ranges.
- Fix one wrong word in burned-in captions without a re-transcribe
One misheard word in a burned-in caption? Re-burn with source_caption_id and a words list instead of running speech-to-text again. Request body included.
- Cover or contain: fit a 16:9 clip into a vertical stack
Sume timeline compose takes video_fit cover, contain or stretch. Work out what each does to a 16:9 clip in the 1080x1920 default stack, then check one frame.
- Cut a clip into 30-second chunks for 30-second post-processing caps
Runway's SDR-to-HDR model and the Magnific upscaler take 30 seconds at most. Cut any clip into 30-second pieces with Sume video trim at $0.02 per cut.
Written by Sume