Whistle keyword biasing vs Sume STT: fixing product names in captions
Whistle can bias decoding toward keywords. Sume STT takes no vocabulary list, so for captions you pass script_text instead. How the two fixes compare.

Whistle lets you inject keywords so that a product name is more likely to be decoded correctly; Sume STT has no such field, so on Sume you fix wording at the caption step with script_text. Both aim at the same problem, a misspelled brand name in a transcript, but they act at different points: Whistle changes the decoding, while script_text replaces the burned wording and keeps the recognised timings.
Cactus Compute lists keyword biasing among Whistle's features in its launch post (read 2026-10-11), alongside word-level timestamps and a 30-second window. The Sume side is from the repository's STT tool contract, which accepts audio_url, language_code, duration_seconds, segmentation and metadata, and from the caption docs.
Where does each fix act?
Biasing is a hint to the recogniser, so it improves the odds but gives no guarantee. Script alignment is deterministic about the text you see, but it needs the script to exist and to resemble what was said.
| Item | Whistle keyword biasing | Sume script_text on a caption job |
|---|---|---|
| Where | Inside local decoding | On the caption render |
| Input | A list of keywords or phrases | The approved script as text |
| Timing source | Whistle word timestamps | STT word timings kept as the source of time |
| Guarantee | Raises the probability of a phrase | Burns your wording, unless alignment fails |
| Failure mode | Name still misspelled | script_alignment_mismatch or script_alignment_failed |
What does Sume's STT step accept?
The STT request has no vocabulary, keyword or dictionary field; the tool description also says diarize and tag_audio_events are fixed server-side and are rejected if sent. What you can set is language_code as a hint, duration_seconds from 1 to 600 to improve the usage reservation, and segmentation of mode sentence for sentence segments. words[] always comes back.
What should you do when alignment fails?
The caption docs say the alignment can fail with script_alignment_mismatch or script_alignment_failed, and recommend simplifying the script or omitting script_text so the STT text is burned. In practice, remove ad-libbed words from the script, or fix the one misheard word by sending words with your own text and times, which skips STT entirely. Send only one of script_text, words, cues and segments.
Can you combine biasing with script_text?
Yes, as two layers. Bias the local decoder so Whistle's own transcript is closer, use that transcript to drive timing and review, and keep the approved script as the final wording in a Sume caption job. The result is a render whose words come from your script and whose times come from recognised speech, so a bad guess in the transcript does not reach the screen.
Keep one habit regardless of the route: compare the burned text with the script before publishing. A caption error on a product name is visible to every viewer, and a mismatch is cheaper to catch in a preview than after upload. For a restyle you do not have to recognise the speech again; send source_caption_id and the new style, and Sume reuses the stored timings.
Which should you choose?
If you run Whistle locally and control the keyword list, bias the decoder and then still check the output. If your audio goes through Sume, keep the approved script and pass it as script_text. When the stakes are high, such as a legal name or a price, read the burned result before publishing, because both methods can still miss.
Sources
Related posts
More in Models
- MiMo V2.6 Pro and Flash in Sume's agent picker: what the catalog says
Sume's model catalog has rows for Xiaomi MiMo V2.6 Pro and Flash behind the OpenRouter switch. What the repo records, and what the API cannot pick.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
Written by Sume