Korean BBQ restaurant reel: korean-ad captions on spoken Hangul
Caption a restaurant reel with Korean speech using Sume's korean-ad style, language ko, and why slam or punch fails with Hangul. Request body and font choices.

To caption a restaurant reel with spoken Korean, send the clip's public HTTPS URL to POST /v1/video-captions with style: "korean-ad" and language: "ko". Sume describes korean-ad as a CapCut-style Hangul karaoke look for Korean speech: one short phrase at a time, in the lower third, with the spoken word changing to a heavy weight.
Do not use slam, punch or tiktok-green for Korean text. Those styles use Latin display faces without Hangul glyphs, and Sume rejects the request with 400 caption_hangul_text_latin_style rather than burn tofu boxes onto your video.
What happens if you send no style
If you omit style, the text of the captions sets it. Korean text resolves to black-outline, which is a safe default: white fill on a thick black outline, in the middle of the frame. It does not resolve to korean-ad, so if you want the ad look with the accent color on the spoken word, say so.
language only tells speech-to-text which language to expect. It never selects the style or the font.
The request
The clip must have audible speech. A silent clip fails with caption_no_speech, and the next action is to use overlay captions. For a menu reel with music only, send authored cues with text, start and end instead, and use a Hangul style for them.
If you have the exact script, pass it as script_text to align the captions to your words rather than to the transcript, which helps with dish names.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: bbq-reel-captions-001" \
-d '{
"video_url": "https://example.com/bbq-reel-clean.mp4",
"style": "korean-ad",
"language": "ko",
"script_text": "오늘 저녁은 숯불 갈비로 정했어요"
}'Pick a font and a look
font is optional and only applies to Hangul styles. Pretendard is the face for korean-ad, and Do Hyeon for black-outline and clip-wipe. Setting font changes the face and keeps the style. design overrides colors, typography, placement, phrasing and motion for one request, for example a brand-red active word.
| Style | Look | Default font |
|---|---|---|
| korean-ad | Ad karaoke, accent on the spoken word | Pretendard |
| black-outline | White fill, thick black outline, middle of frame | Do Hyeon |
| clip-wipe | Each word wipes in left to right, clear at small sizes | Do Hyeon |
| pill-karaoke | Card in a dark pill, color changes with the voice | Pretendard |
| slam, punch, tiktok-green | Latin faces, Hangul rejected | Not for Korean |
Where the speech comes from
If your reel has no voice, Sume's avatar talking-video route can produce a spoken Korean-script clip with inline captions, and a Korean script on a Latin style is rejected there too with the same error. Check the final frames on a phone at arm's length: menu names should be readable without zooming.
Sources
Related posts
More in Use cases
- Landscape listing photos in a vertical reel: cover, contain or blur?
A 3:2 photo in a 9:16 reel shows only 37.5% of its width with fit cover. Pick the fit per photo in Timeline 1.0: eight photos and a bed cost $0.225 on Sume.
- Live stream avatar: queue pre-rendered reaction clips instead
For a stream overlay, render ten short Sume avatar clips ahead of time and trigger them from chat events. Ten 6-second clips cost $14.70 on plus.
- Lock the look with one image, then animate it on three video models
Make one approved image, then use it as first_frame on wan-3.0, kling-3 and minimax-h3-max. Same start, three motions, one prompt, so you can compare.
- Lofi rain-on-window loop with sound: Gemini Omni 10 seconds
Make a 10-second lofi rainy window clip with Gemini Omni on Sume. Ambient audio wording, one image as first and last frame for a loop, and the cost to draft.
Written by Sume