Whistle 16.9 MB on-device model vs Sume STT: which clips go where
Cactus Whistle transcribes 30-second clips on the device; Sume STT bills $0.01 per audio minute in the cloud. Which audio belongs on which, with a table.

Use Cactus Whistle when the audio is a short clip in one of its seven languages and must never leave the device. Use Sume STT when the audio is longer, in another language, or already sits in a cloud workflow, because Sume transcribes a file you point it at for $0.01 per audio minute. Both return word timings, so the choice is about where the audio is allowed to go, not about output shape.
Whistle is a 16.9 MB speech recognition model from Cactus Compute. In the company's launch post (read 2026-10-11) it takes 16 kHz mono audio, handles up to 30 seconds per pass, and covers English, German, French, Spanish, Italian, Dutch and Polish. Cactus lists a time to first token of 11.1 ms on an Apple M4 Pro CPU and says it runs entirely on-device with no dependencies. Sume does not ship Whistle; this post is about the routing decision between a local model and a hosted job.
What does each option accept?
The limits differ in kind. Whistle's limit is a window per pass. Sume STT's limit is a file and a per-minute price. The table lines up the facts each vendor states.
| Question | Cactus Whistle | Sume STT 1.0 |
|---|---|---|
| Where it runs | On the device, CPU, audio stays local | Sume cloud job |
| Input | 16 kHz mono audio | audio_url, a public HTTPS URL (Sume media host preferred) |
| Length per pass | Up to 30 seconds | Catalog describes a maximum of 10 minutes per request |
| Languages | English, German, French, Spanish, Italian, Dutch, Polish | Auto-detect, or a language_code hint such as ko |
| Word timings | Start, end and probability per word | words[] with word, start, end |
| Price | Model weights run on your hardware | $0.01 per audio minute |
When should a clip stay on the device?
Keep it local when the recording is privacy-bound, when the connection is unreliable, or when you need a transcript in milliseconds inside an app. A push-to-talk note, a wake-phrase check or a short voice command all fit inside 30 seconds and one of the seven languages.
Cactus also reports that Whistle beats Whisper base on word error rate for several benchmarks (LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22, FLEURS average) while Whisper base is better on TED-LIUM, AMI and the MLS average. So the quality claim is dataset-specific. Test on your own audio before you commit.
When should you send it to Sume STT?
Send it to Sume when the clip is longer than one window, when the language is outside Whistle's list, or when the transcript feeds something else Sume does. Sume STT 1.0 is POST /v1/stt-1.0/transcribe; you submit an audio_url, an Idempotency-Key and an optional language_code or duration_seconds, then read the job result. words[] always comes back, and segmentation with mode sentence adds sentence segments.
For a clip that is still inside a video, video inspect with transcribe true runs the same STT at the same $0.01 per audio minute, and audio detach pulls the track out as wav or mp3 first. Jobs are asynchronous, so store the job id and poll its status rather than resubmitting a paid request.
How do you decide in practice?
Ask three questions in order.
- Is the audio allowed to leave the device? If not, Whistle or another local model is the only answer.
- Does it fit 30 seconds and one of the seven languages? If yes, local is the cheapest and fastest path.
- Otherwise, upload the file, pay per minute, and take the word timings into captions or a timeline on Sume.
Sources
Related posts
More in Models
- Whistle's 30-second window: passes for 5 minutes vs Sume STT
A 5-minute recording is ten 30-second Whistle passes on the device, or one Sume STT job at $0.05. How to plan chunking and when to skip it.
- Whistle keyword biasing vs Sume STT: fixing product names in captions
Whistle can bias decoding toward keywords. Sume STT takes no vocabulary list, so for captions you pass script_text instead. How the two fixes compare.
- MiMo V2.6 Pro and Flash in Sume's agent picker: what the catalog says
Sume's model catalog has rows for Xiaomi MiMo V2.6 Pro and Flash behind the OpenRouter switch. What the repo records, and what the API cannot pick.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
Written by Sume