Korean clip? Whistle's seven languages vs Sume STT
Cactus Whistle lists English, German, French, Spanish, Italian, Dutch and Polish. For a Korean clip use Sume STT with a ko hint, then caption it.

Cactus Whistle will not transcribe a Korean clip, because its published language list is English, German, French, Spanish, Italian, Dutch and Polish. For that language, send the file to Sume STT with language_code set to ko (or omit it and let auto-detect choose), and take the word timings into a caption job.
The language list is from Cactus Compute's Whistle launch post (read 2026-10-11), which also says the model detects the language automatically within that set. Detecting automatically inside seven languages is not the same as recognizing any language, so a Korean clip may come back as a wrong guess in one of the seven rather than as an error. Check the language before you trust the text.
Which languages are covered where?
Sume documents the hint rather than a language list. In the video inspect docs, language_code is an STT hint such as en or ko, and without it STT uses auto-detect. The caption docs use the same idea: language tells speech-to-text which language to expect and never selects the style or the font.
| Item | Cactus Whistle | Sume STT and captions |
|---|---|---|
| Stated languages | English, German, French, Spanish, Italian, Dutch, Polish | No fixed list on the docs pages read; hint with language_code, for example ko |
| Unspecified language | Automatic detection within its set | Auto-detect when language_code is omitted |
| Korean speech | Not in the list | Hint ko; use a Hangul caption style |
What does a Korean caption job need?
The caption docs separate language from look. A Korean clip should use a Hangul-capable style: black-outline is the default for Korean text, and korean-ad gives the one-phrase-at-a-time karaoke look when you set language to ko. Sending Korean text to slam, punch or tiktok-green is refused with caption_hangul_text_latin_style, because those Latin display faces would render tofu boxes.
A standalone caption job reserves and captures $0.20 for a video up to 60 seconds under the current estimate; check GET /v1/catalog for the live price.
How do you keep the wording right?
Speech recognition can misspell product names in any language. If you have the approved script, pass it as script_text on the caption job: Sume keeps the STT word timings as the source of time and aligns the burned text to your wording. The alignment can fail with script_alignment_mismatch or script_alignment_failed, and the docs recommend simplifying the script or leaving script_text out in that case.
Can you mix both routes?
Yes, and a mixed pipeline is possible. A product with English and Korean users can run Whistle on the device for English voice notes and send Korean recordings to Sume. Keep the routing decision in your app: read the user's locale or a first-pass language guess, and only upload when the clip is outside the local model's set. Record which route produced each transcript, so a wrong guess can be traced.
Also keep the unit of work in mind. Whistle's window is 30 seconds per pass, while Sume STT takes the file as one job. A 45-second Korean ad read therefore needs no chunking on the Sume side.
What is the simplest rule?
If the language is on Whistle's list, the clip is under 30 seconds and the audio has to stay local, use Whistle. If the language is not on the list, hint the language and use Sume STT or the caption job directly on the video.
Sources
Related posts
More in Models
- Muse Spark 1.3 on Sume: picker row, tool use and media jobs
Meta says Muse Spark 1.3 uses about 20% fewer tool calls. Sume lists it as a catalog row behind the OpenRouter switch; the API cannot pick it.
- Nova 2.5 Sonic for a narration file? Speech-to-speech vs Sume TTS
Nova 2.5 Sonic is built for live voice agents. For a finished narration file, Sume TTS takes a transcript up to 20,000 characters and returns audio as a job.
- Qwen-Image-2.1-Turbo runs 8 steps; what does Sume's Qwen row expose?
Steps, CFG and seed are model-card settings. Sume's qwen/qwen-image row lists none of them: only ratio, n 1-4, references, output format. Check with one GET.
- Reka Rho-1: can you generate video with it today?
Reka Rho-1 is a 19B research preview announced October 5, 2026, with no public weights, API or demo link. For video you can run today, check Sume's catalog.
Written by Sume