Sume caption design overrides: one key, one change, bad value 400s
Sume video captions design overrides merge over the style, so one key changes one thing. A bad value fails with 400 at request time, before any render.

On Sume's standalone captions endpoint, design overrides merge field by field over the chosen style, so sending design.colors.active changes the emphasis color and nothing else. A number outside its documented range returns 400 at request time, so a wrong look does not cost you a render. punch and tiktok-green do not support design. Source: Video captions, read 2026-10-06.
Which groups can I override?
Every field is optional and the five groups are the whole surface.
| Group | Fields |
|---|---|
colors | base, active, stroke, accent, accent_deep, card |
typography | base_weight, active_weight, active_scale, font_size_ratio, safe_width_ratio, stroke_width_px |
placement | anchor_ratio, landscape_anchor_ratio |
phrasing | max_words, max_chars, pause_seconds |
motion | enter_seconds, exit_seconds, emphasis_in_seconds, emphasis_out_seconds |
What values are accepted?
Colors are hex, rgb(), rgba() or transparent. Other CSS syntax is rejected and never reaches the render document. card: null shows no card. Read the documented range of each field before you wire a slider to the API.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-cyan-001" \
-d '{
"video_url": "https://example.com/clean.mp4",
"style": "black-outline",
"design": { "colors": { "active": "#22D3EE" } }
}'How do I restyle without paying for transcription again?
Send source_caption_id from an earlier caption job with a new style or design, and Sume reuses that transcript instead of running speech-to-text again.
Sources
Related posts
More in Media tools
- Korean caption fonts on Sume: why weight-shift needs Pretendard
Sume's weight-shift and korean-ad caption styles animate a font weight axis that only Pretendard has. Other faces keep color and scale emphasis only.
- Captions for a voiceover: reuse TTS word timings or run STT?
If you generated the voice on Sume, ask for timestamps.words and skip a 1-cent-a-minute STT pass. If the audio is someone else's, run STT. Here is the split.
- Captions language field: a speech hint, not a style or font
On Sume video captions, language only hints speech-to-text. It never picks the style or font, so Korean needs a Hangul style such as black-outline.
- Check a Short has sound and picture before adding it to a YouTube show
YouTube asks shows for high-quality video and sound and says no-visuals videos likely will not count. Probe each Short with Sume video inspect first.
Written by Sume