Pick a caption style from the STT language_code before you burn

Run Sume STT on the detached audio, read language_code, and choose korean-ad or slam before POST /v1/video-captions. Avoids the Hangul-on-Latin 400.

5 min readSume
All posts

If you want the caption style to follow the language of the speech, read language_code from the Sume STT result and map it to a style before you call POST /v1/video-captions. Korean speech goes to a Hangul style such as korean-ad. Latin-script speech goes to slam, punch or tiktok-green. The mapping matters because the Latin display faces have no Hangul glyphs, and Sume answers a Korean text on one of them with 400 caption_hangul_text_latin_style instead of burning boxes.

The chain takes three steps: detach the audio, transcribe it, burn the captions. Each step is a separate Sume job, and none of them needs a file outside media.sume.com.

Where the language comes from

The STT 1.0 result carries text, words[], and, when the engine reports them, language_code and language_probability. Send language_code in the request only as a hint. If you omit it, STT uses auto-detect, and the detected code comes back in the result. The audio detach page names the input shape for this step: format: wav, sample_rate: 16000, channels: mono.

The STT route is POST /v1/stt-1.0/transcribe, listed on the timeline audio page. Over hosted MCP the same step is stt_create, which is on the paid-generation list in MCP tools and gates.

The mapping

The docs only describe Latin faces and Hangul faces. They do not document other scripts. If your STT result returns a code for a script outside those two, test one clip before you batch.

Language code to caption style, from the Sume video-captions and tools docs (read 2026-10-05)
language_code starts withStyle to sendWhy
kokorean-ad, or black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasisHangul faces. The docs recommend a Hangul style for Korean speech.
en and other Latin-script codesslam (default), punch, tiktok-greenLatin display faces. Latin text on slam works as before.
null or low language_probabilitySend no style, or review firstWith no style, the caption text decides: Latin gets slam, Korean gets black-outline.

A small mapper

This pure function has no network call. Feed it the language_code from the finished STT job, and put the result in the caption request body.

def caption_style(language_code):
    code = (language_code or "").lower()
    if code.startswith("ko"):
        return "korean-ad"
    return "slam"

body = {
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": caption_style("ko"),
    "language": "ko",
}
print(body)

Details that trip people

language on the caption request is a speech-to-text hint. It never selects the style or the font. That is why you set style yourself. korean-ad is not what you get by default for Korean: with no style, Korean text resolves to black-outline.

If you already have the transcript from STT, you do not need a second transcription inside the caption job. Send words or cues with text, start and end, and Sume burns your text at those times without running speech-to-text again. You can send only one of script_text, words, cues and segments.

Standalone captions cost $0.20 per job for a video of at most 60 seconds, under the current fixed estimate. The STT step is billed per audio minute and the detach step per job. Check the live rates with GET /v1/catalog before you budget a batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume