Whistle keyword biasing vs Sume STT: fixing product names in captions

Whistle can bias decoding toward keywords. Sume STT takes no vocabulary list, so for captions you pass script_text instead. How the two fixes compare.

4 min readSume
All posts

Whistle lets you inject keywords so that a product name is more likely to be decoded correctly; Sume STT has no such field, so on Sume you fix wording at the caption step with script_text. Both aim at the same problem, a misspelled brand name in a transcript, but they act at different points: Whistle changes the decoding, while script_text replaces the burned wording and keeps the recognised timings.

Cactus Compute lists keyword biasing among Whistle's features in its launch post (read 2026-10-11), alongside word-level timestamps and a 30-second window. The Sume side is from the repository's STT tool contract, which accepts audio_url, language_code, duration_seconds, segmentation and metadata, and from the caption docs.

Where does each fix act?

Biasing is a hint to the recogniser, so it improves the odds but gives no guarantee. Script alignment is deterministic about the text you see, but it needs the script to exist and to resemble what was said.

Two ways to fix a misheard product name (Cactus blog and Sume docs, read 2026-10-11)
ItemWhistle keyword biasingSume script_text on a caption job
WhereInside local decodingOn the caption render
InputA list of keywords or phrasesThe approved script as text
Timing sourceWhistle word timestampsSTT word timings kept as the source of time
GuaranteeRaises the probability of a phraseBurns your wording, unless alignment fails
Failure modeName still misspelledscript_alignment_mismatch or script_alignment_failed

What does Sume's STT step accept?

The STT request has no vocabulary, keyword or dictionary field; the tool description also says diarize and tag_audio_events are fixed server-side and are rejected if sent. What you can set is language_code as a hint, duration_seconds from 1 to 600 to improve the usage reservation, and segmentation of mode sentence for sentence segments. words[] always comes back.

What should you do when alignment fails?

The caption docs say the alignment can fail with script_alignment_mismatch or script_alignment_failed, and recommend simplifying the script or omitting script_text so the STT text is burned. In practice, remove ad-libbed words from the script, or fix the one misheard word by sending words with your own text and times, which skips STT entirely. Send only one of script_text, words, cues and segments.

Can you combine biasing with script_text?

Yes, as two layers. Bias the local decoder so Whistle's own transcript is closer, use that transcript to drive timing and review, and keep the approved script as the final wording in a Sume caption job. The result is a render whose words come from your script and whose times come from recognised speech, so a bad guess in the transcript does not reach the screen.

Keep one habit regardless of the route: compare the burned text with the script before publishing. A caption error on a product name is visible to every viewer, and a mismatch is cheaper to catch in a preview than after upload. For a restyle you do not have to recognise the speech again; send source_caption_id and the new style, and Sume reuses the stored timings.

Which should you choose?

If you run Whistle locally and control the keyword list, bias the decoder and then still check the output. If your audio goes through Sume, keep the approved script and pass it as script_text. When the stakes are high, such as a legal name or a price, read the burned result before publishing, because both methods can still miss.

Sources

Related posts

More in Models

All Models posts

Written by Sume