Korean karaoke captions: korean-ad and language ko on Sume
Burn Korean karaoke-style captions with style korean-ad and language ko on /v1/video-captions: one phrase at a time, the spoken word in a heavier weight.

How do you add Korean karaoke captions to a video through an API? Send the clip to POST /v1/video-captions with style: "korean-ad" and language: "ko". The docs describe korean-ad as CapCut-style Hangul karaoke for Korean speech: one short phrase at a time, lower third, with the spoken word shifting to a heavy weight.
Do not use slam, punch or tiktok-green for Korean. They draw in Latin display faces with no Hangul glyphs, and Sume rejects Korean text on them with caption_hangul_text_latin_style rather than rendering tofu boxes.
What the style does
korean-ad is weight-shift with an accent colour on the spoken word. It is not what an omitted style resolves to: an omitted style gives black-outline for Korean, so name korean-ad when you want the ad treatment. weight-shift and korean-ad animate the wght axis, which only Pretendard carries, so a static face keeps colour and scale emphasis but loses the weight travel.
Choosing among the Hangul identities
| Style | Look |
|---|---|
black-outline | White fill, thick black outline, mid-frame; the safe default |
weight-shift | Phrase cards; the spoken word takes the weight |
highlight | An accent block sweeps in behind the spoken word |
pill-karaoke | Card in a dark pill; colour tracks the voice |
clip-wipe | Each word wiped in left to right; clear at small sizes |
editorial-emphasis | Two-line left-aligned card with a large phrase-final word |
Send the request
language only tells speech-to-text what to expect; it never selects the style or the font. Add script_text if you want the burned wording to follow a script, with timings still coming from speech-to-text.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ko-karaoke-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "korean-ad",
"language": "ko",
"design": { "colors": { "active": "#22D3EE" } }
}'Restyle without a second transcript
To try a different look, send source_caption_id instead of video_url. Sume reuses the source video and the word timings it already has, so no second speech-to-text runs, and billing is still a render at $0.20 for videos up to 60 seconds. If the script and the audio disagree, alignment can fail with script_alignment_mismatch; omit script_text to burn the transcribed wording.
Sources
Related posts
More in Developers
- Workers KV jurisdictions: keep Sume job records in region
Cloudflare made Workers KV jurisdictions generally available on Oct 2, 2026. Here is how to key Sume job_id records into a region-scoped namespace.
- LangGraph custom image: keep SUME_API_KEY out of it
langgraph-cli 0.4.32 adds an --image-uri flag for self-hosted custom containers. Inject SUME_API_KEY as a runtime env var; never bake it into the image.
- Luma callback_url or polling: which Sume job mode matches
Luma's API docs list keyframes, loop and callback_url for ray-2. On Sume the equivalent choice is job mode: async, sync up to 30 s, subscribe or webhook.
- Luma API callback_url vs Sume callback_url: signing and retries
Luma's video API takes a callback_url, and so does Sume's /v1/videos. What Sume's callback is signed with, how often it retries, and a Python verifier.
Written by Sume