Burned-in captions: ALL CAPS or as written? What each Sume style does
Sume slam, punch and tiktok-green upper-case every Latin word at render time; Hangul-identity styles burn text as written. What script_text and cues change.

On Sume, the three Latin caption styles (slam, punch, tiktok-green) always burn ALL CAPS: the renderer upper-cases each word, so the case in your script_text, words or STT transcript does not survive. The Hangul-identity styles burn text exactly as written, with no case folding. Each standalone caption job costs $0.20 for a video up to 60 seconds.
Read on 2026-10-03: the Sume video-captions docs page and the caption renderer source it documents.
Where does the capitalization happen?
In slam, the word timeline is built with toUpperCase() before render, and the template also sets text-transform: uppercase on every word. punch and tiktok-green draw through a Remotion composition that sets textTransform: "uppercase" on the whole caption. So on those three styles, mixed-case input is a no-op for the look.
| Style group | Face | Latin case in the burn |
|---|---|---|
| slam | Anton | Upper-cased |
| punch, tiktok-green | TheBoldFont | Upper-cased |
| black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasis, korean-ad | Hangul faces (Pretendard, Do Hyeon and others) | Rendered as written, no case transform |
What do script_text, words and cues change?
They change the wording, not the case rule. script_text keeps STT timings and aligns your wording to them; words (word-level) and cues or segments (phrase cards) skip speech-to-text and burn your copy at your times. They are mutually exclusive. On a Latin style all of them still come out upper-case. On a Hangul identity they come out exactly as you typed them, including mixed-case product names or model numbers.
A silent clip has no speech to transcribe, so it fails as caption_no_speech unless you pass cues.
How do I get sentence case or Title Case?
Name a Hangul-identity style with English text. The request guard only refuses the opposite pairing, Hangul wording on a Latin style (caption_hangul_text_latin_style), so Latin text on weight-shift is accepted and not case-folded. The catch: those styles are authored for Korean, and I did not verify in code how well their faces draw English letters. Burn one test clip ($0.20) and look before you batch.
If you need a specific Latin display face in mixed case, the styles above do not offer it. Their font option accepts Hangul faces only, and naming one next to a Latin style returns caption_font_requires_hangul_style.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-as-written-001" \
-d '{
"video_url": "https://example.com/clip.mp4",
"style": "weight-shift",
"script_text": "Meet the new iPhone case, in three colors."
}'Does the default style change anything?
Omitting style resolves to slam for Latin wording, so a default English job is ALL CAPS. It resolves to black-outline when the wording is mostly Hangul.
Sources
Related posts
More in Media tools
- Can a 16:9 video be an Instagram Reel? The range says yes
Instagram accepts Reels from 1.91:1 to 9:16, so 16:9 (1.78) is inside the range. How to render a 1920x1080 cut or a vertical one from the same Sume clips.
- Caption a Recast video: burn captions after the person swap
Run h3-max-recast first, then send its output to POST /v1/video-captions. The kept source audio gives the captions speech to transcribe.
- Caption micro-drama dialogue per shot: $0.20 per clip up to 60 s
Video captions cost $0.20 per job for clips up to 60 seconds. Caption each shot before assembly; a silent clip fails with caption_no_speech. Budget math.
- Caption line breaks: Netflix bottom-heavy pyramid rule vs Sume cards
Netflix wants two-line captions broken as a bottom-heavy pyramid. Sume caption cards break by phrasing settings. How to approximate the rule and where it ends.
Written by Sume