Korean burned-in captions in a CapCut style: Sume Hangul styles
Sume burns Korean captions with six Hangul styles and 29 open-license fonts at $0.20 per clip up to 60 seconds. Latin styles reject Korean with a 400.

For Korean speech burned into a video, Sume ships six Hangul caption styles and 29 fonts through one API call, $0.20 per clip up to 60 seconds. The default for Korean text is black-outline, which the docs describe as CapCut-style white text on a thick black outline in the middle of the frame. I could not read CapCut's own pricing page (the request returned a 404), so this post makes no claim about CapCut features or cost; it only describes what Sume does.
The main trap is the style/language mismatch. Latin styles slam, punch and tiktok-green have no Hangul glyphs, and Sume rejects Korean text sent to them with a 400 instead of rendering empty boxes.
| Style | Look |
|---|---|
| black-outline | White fill on a thick black outline, middle of frame; the default for Korean |
| weight-shift | Phrase cards; the spoken word takes the heavy weight |
| highlight | An accent block moves in behind the spoken word |
| pill-karaoke | Card in a dark pill; color changes with the voice |
| clip-wipe | Each word enters with a left-to-right wipe; clearest at small phone sizes |
| editorial-emphasis | Left-aligned two-line card; last word of the phrase drops to a second, larger line |
Karaoke ad look and fonts
korean-ad is a separate ad-karaoke style: one short phrase at a time in the lower third, with the spoken word at a heavy weight. It is not what you get by default, so select it and set language to ko. Fonts apply to Hangul styles only: pretendard is the baseline, do-hyeon is the thick rounded classic, black-han-sans and gasoek-one are impact faces, and gmarket-sans is the retail display staple.
All 29 fonts use the SIL Open Font License 1.1 and are bundled in the renderer. Names outside the list are rejected, and there is no silent fallback face. Weight travel on weight-shift and korean-ad works only with Pretendard, the one font with a weight axis.
Control and correction
Add design overrides for one request: colors, typography, placement, phrasing and motion. Numbers outside their documented range return a 400 at request time, so a bad look fails before you pay. Send script_text if you want the exact spelling of product names: Sume keeps the speech-recognition timings and aligns your text to them, and typed errors such as script_alignment_mismatch tell you when it cannot.
To try another style, send source_caption_id with a new style. Sume reuses the stored word timings, so speech recognition does not run again, and the $0.20 price is unchanged.
- Send language: "ko" as a hint to speech-to-text; it never selects the style or the font.
- Silent clips fail with caption_no_speech; send cues with text, start and end instead.
- Video must be a public HTTPS URL; private or signed URLs are rejected.
Verdict
If your captions are made in a phone editor by one person, use that editor. If you need the same Korean look on a hundred ads, produced from code with exact spelling control, the six styles and the restyle path are the reason to call an API.
Sources
Related posts
More in Comparisons
- Within 30 Elo of Wan 3.0 for less: MiniMax H3 and H3 Max on Sume
MiniMax H3 (1,137) and H3 Max (1,130) sit 19 and 26 Elo behind Wan 3.0 on the AA board. Sume bills $4.50 and $6.00 a minute at 768p; Wan is $7.50 at 720p.
- Zoom deepfake risk detection: what it means for avatar video
Zoom lists real-time alerts for synthetic audio or video next to meeting avatars. What the page says, what it leaves open, and how to label a Sume clip.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
Written by Sume