Pick a caption style from the STT language_code before you burn
Run Sume STT on the detached audio, read language_code, and choose korean-ad or slam before POST /v1/video-captions. Avoids the Hangul-on-Latin 400.

If you want the caption style to follow the language of the speech, read language_code from the Sume STT result and map it to a style before you call POST /v1/video-captions. Korean speech goes to a Hangul style such as korean-ad. Latin-script speech goes to slam, punch or tiktok-green. The mapping matters because the Latin display faces have no Hangul glyphs, and Sume answers a Korean text on one of them with 400 caption_hangul_text_latin_style instead of burning boxes.
The chain takes three steps: detach the audio, transcribe it, burn the captions. Each step is a separate Sume job, and none of them needs a file outside media.sume.com.
Where the language comes from
The STT 1.0 result carries text, words[], and, when the engine reports them, language_code and language_probability. Send language_code in the request only as a hint. If you omit it, STT uses auto-detect, and the detected code comes back in the result. The audio detach page names the input shape for this step: format: wav, sample_rate: 16000, channels: mono.
The STT route is POST /v1/stt-1.0/transcribe, listed on the timeline audio page. Over hosted MCP the same step is stt_create, which is on the paid-generation list in MCP tools and gates.
The mapping
The docs only describe Latin faces and Hangul faces. They do not document other scripts. If your STT result returns a code for a script outside those two, test one clip before you batch.
| language_code starts with | Style to send | Why |
|---|---|---|
| ko | korean-ad, or black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasis | Hangul faces. The docs recommend a Hangul style for Korean speech. |
| en and other Latin-script codes | slam (default), punch, tiktok-green | Latin display faces. Latin text on slam works as before. |
| null or low language_probability | Send no style, or review first | With no style, the caption text decides: Latin gets slam, Korean gets black-outline. |
A small mapper
This pure function has no network call. Feed it the language_code from the finished STT job, and put the result in the caption request body.
def caption_style(language_code):
code = (language_code or "").lower()
if code.startswith("ko"):
return "korean-ad"
return "slam"
body = {
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": caption_style("ko"),
"language": "ko",
}
print(body)Details that trip people
language on the caption request is a speech-to-text hint. It never selects the style or the font. That is why you set style yourself. korean-ad is not what you get by default for Korean: with no style, Korean text resolves to black-outline.
If you already have the transcript from STT, you do not need a second transcription inside the caption job. Send words or cues with text, start and end, and Sume burns your text at those times without running speech-to-text again. You can send only one of script_text, words, cues and segments.
Standalone captions cost $0.20 per job for a video of at most 60 seconds, under the current fixed estimate. The STT step is billed per audio minute and the detach step per job. Check the live rates with GET /v1/catalog before you budget a batch.
Sources
Related posts
More in Developers
- STT with no punctuation: Sume still cuts sentence segments on silence
Sume STT's sentence segmentation splits on terminal punctuation, and on silence when a run has none. What boundary_lead_ms 70 does and a Python call.
- STT words into caption words: rename word to text, stay under 60 s
Sume STT returns words as {word, start, end}; the caption job wants {text, start, end}, end above start, 60 s or less, 1200 words at most. Map and filter.
- Submit Sume jobs from Go with a bounded pool and idempotency keys
Bound a Go worker pool to your workspace in-flight headroom, send a stable Idempotency-Key per item, and stop treating a full queue as a crash. Stdlib only.
- Sume 400 unsupported_capability on 4:5: fall back to 3:4 in Python
A 4:5 request on a Sume video model returns 400 unsupported_capability before billing. Catch it in Python, resubmit at 3:4, then crop to Meta's 4:5 Feed ratio.
Written by Sume