Reel captions in two languages: language hint vs auto-detect

How Sume video captions pick a language and a style: the optional language hint, auto-detect, the Hangul rule that returns 400, and a Korean Reel example.

6 min readSume
All posts

Short answer: hint when you know, omit when you do not

For a Reel whose speech is in one language, send language (ko, en and so on) to video captions. It is a speech-to-text hint. If you leave it out, speech-to-text detects the language itself. For a Reel that switches languages mid-sentence the docs say nothing about code-switching behavior, so we do not promise a result; the safe design is to test a short clip, or split the clip by language and caption each part.

The field that people expect to do more than it does is language. The docs are explicit: style selects the look and the motion, design changes that look, font selects the face of the text, and language only tells speech-to-text which language to expect. It never selects the style or the font.

The Hangul rule that can fail a request

Caption styles slam, punch and tiktok-green use Latin display faces with no Hangul glyphs. If Korean text reaches one of them, the API returns 400 with caption_hangul_text_latin_style instead of burning boxes into your video and charging you for it. The same applies to a Hangul font on a Latin style (caption_font_requires_hangul_style). If you send no style, text decides: Latin text resolves to slam, Korean text to black-outline.

So a bilingual account that posts both English and Korean Reels should not hard-code one style for every request. Either omit style and let the text decide, or choose the style from the language you already know. For Korean ad-style karaoke, korean-ad is the named look, and it is not what you get by default; the default is black-outline.

Decision table

Use the row that matches your clip. Each behavior is from the video captions page.

Caption language and style behavior from the Sume docs (read 2026-10-06)
Your clipWhat to sendWhat Sume does
English speechlanguage en (optional), no styleBurns the STT text in slam, the Latin default
Korean speechlanguage ko, style black-outline or korean-adBurns Hangul in a Hangul-capable face
Korean text, style slamDo not400 caption_hangul_text_latin_style
Unknown languageOmit languageSTT detects the language
Silent clipcues or segments with text, start, endNo STT; burns your text; speech mode fails with caption_no_speech
Fix a wrong wordwords with source_caption_idReuses the source video and timings; no second STT

A Korean Reel, with placement and phrasing

The request below burns Korean speech with the safe default style, a hint, a lower caption line and short phrases. anchor_ratio is the center of the line as a fraction of the frame height, and max_words limits words per phrase; both are design overrides, applied on top of the style for this request only. Out-of-range values return 400, so a bad look fails before you are charged.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: reel-captions-ko-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/reel-ko.mp4",
    "language": "ko",
    "style": "black-outline",
    "design": {"placement": {"anchor_ratio": 0.62},
               "phrasing": {"max_words": 4}}
  }'

Price and where the captions end up

Each accepted standalone caption job reserves and captures $0.20 of usage, quoted for videos up to 60 seconds under the current fixed estimate; check GET /v1/catalog for the live number, and for longer Reels do not assume the same price. The result is a captioned video URL on media.sume.com, and the input must be a public HTTPS video URL that Sume can fetch. Poll the job envelope with GET /v1/jobs/:id/status and read the result when it is ready.

If you need a clip's language before you caption it, video inspect can transcribe with an optional language_code hint at $0.01 per audio minute, and a frames: false call returns only the probe, which Sume bills by its Modal compute rather than as a transcript. For a bilingual account, the common mistake is to caption first and discover the language afterwards. Use a probe-only inspect to learn has_audio before you pay for captions.

A practical workflow for an account that posts in both languages: keep one function that takes the clip and a language code, sets language from that code, leaves style unset for English so the Latin default applies, and passes black-outline or korean-ad for Korean. Give every call a stable idempotency key made from the clip id and the language, so a retry replays the same job instead of billing the clip twice. If the look is wrong after the first render, do not re-run speech-to-text: send source_caption_id with a new style and Sume reuses the source video and the word timings it already has, at the same price as any render.

For Korean text you can also choose a font such as pretendard, a clean baseline, or do-hyeon, the thick rounded face described as the CapCut classic. A Hangul font on a Latin style is rejected, so the font choice has to follow the style choice. Place the caption line with anchor_ratio, and check it in the app's own preview before you publish, because the docs do not give safe-zone numbers for any platform.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume