Korean clip? Whistle's seven languages vs Sume STT

Cactus Whistle lists English, German, French, Spanish, Italian, Dutch and Polish. For a Korean clip use Sume STT with a ko hint, then caption it.

4 min readSume
All posts

Cactus Whistle will not transcribe a Korean clip, because its published language list is English, German, French, Spanish, Italian, Dutch and Polish. For that language, send the file to Sume STT with language_code set to ko (or omit it and let auto-detect choose), and take the word timings into a caption job.

The language list is from Cactus Compute's Whistle launch post (read 2026-10-11), which also says the model detects the language automatically within that set. Detecting automatically inside seven languages is not the same as recognizing any language, so a Korean clip may come back as a wrong guess in one of the seven rather than as an error. Check the language before you trust the text.

Which languages are covered where?

Sume documents the hint rather than a language list. In the video inspect docs, language_code is an STT hint such as en or ko, and without it STT uses auto-detect. The caption docs use the same idea: language tells speech-to-text which language to expect and never selects the style or the font.

Language handling, Whistle vs Sume STT (Cactus blog and Sume docs, read 2026-10-11)
ItemCactus WhistleSume STT and captions
Stated languagesEnglish, German, French, Spanish, Italian, Dutch, PolishNo fixed list on the docs pages read; hint with language_code, for example ko
Unspecified languageAutomatic detection within its setAuto-detect when language_code is omitted
Korean speechNot in the listHint ko; use a Hangul caption style

What does a Korean caption job need?

The caption docs separate language from look. A Korean clip should use a Hangul-capable style: black-outline is the default for Korean text, and korean-ad gives the one-phrase-at-a-time karaoke look when you set language to ko. Sending Korean text to slam, punch or tiktok-green is refused with caption_hangul_text_latin_style, because those Latin display faces would render tofu boxes.

A standalone caption job reserves and captures $0.20 for a video up to 60 seconds under the current estimate; check GET /v1/catalog for the live price.

How do you keep the wording right?

Speech recognition can misspell product names in any language. If you have the approved script, pass it as script_text on the caption job: Sume keeps the STT word timings as the source of time and aligns the burned text to your wording. The alignment can fail with script_alignment_mismatch or script_alignment_failed, and the docs recommend simplifying the script or leaving script_text out in that case.

Can you mix both routes?

Yes, and a mixed pipeline is possible. A product with English and Korean users can run Whistle on the device for English voice notes and send Korean recordings to Sume. Keep the routing decision in your app: read the user's locale or a first-pass language guess, and only upload when the clip is outside the local model's set. Record which route produced each transcript, so a wrong guess can be traced.

Also keep the unit of work in mind. Whistle's window is 30 seconds per pass, while Sume STT takes the file as one job. A 45-second Korean ad read therefore needs no chunking on the Sume side.

What is the simplest rule?

If the language is on Whistle's list, the clip is under 30 seconds and the audio has to stay local, use Whistle. If the language is not on the list, hint the language and use Sume STT or the caption job directly on the video.

Sources

Related posts

More in Models

All Models posts

Written by Sume