Whistle 16.9 MB on-device model vs Sume STT: which clips go where

Cactus Whistle transcribes 30-second clips on the device; Sume STT bills $0.01 per audio minute in the cloud. Which audio belongs on which, with a table.

5 min readSume
All posts

Use Cactus Whistle when the audio is a short clip in one of its seven languages and must never leave the device. Use Sume STT when the audio is longer, in another language, or already sits in a cloud workflow, because Sume transcribes a file you point it at for $0.01 per audio minute. Both return word timings, so the choice is about where the audio is allowed to go, not about output shape.

Whistle is a 16.9 MB speech recognition model from Cactus Compute. In the company's launch post (read 2026-10-11) it takes 16 kHz mono audio, handles up to 30 seconds per pass, and covers English, German, French, Spanish, Italian, Dutch and Polish. Cactus lists a time to first token of 11.1 ms on an Apple M4 Pro CPU and says it runs entirely on-device with no dependencies. Sume does not ship Whistle; this post is about the routing decision between a local model and a hosted job.

What does each option accept?

The limits differ in kind. Whistle's limit is a window per pass. Sume STT's limit is a file and a per-minute price. The table lines up the facts each vendor states.

Whistle and Sume STT 1.0 side by side (Cactus blog and Sume repo and docs, read 2026-10-11)
QuestionCactus WhistleSume STT 1.0
Where it runsOn the device, CPU, audio stays localSume cloud job
Input16 kHz mono audioaudio_url, a public HTTPS URL (Sume media host preferred)
Length per passUp to 30 secondsCatalog describes a maximum of 10 minutes per request
LanguagesEnglish, German, French, Spanish, Italian, Dutch, PolishAuto-detect, or a language_code hint such as ko
Word timingsStart, end and probability per wordwords[] with word, start, end
PriceModel weights run on your hardware$0.01 per audio minute

When should a clip stay on the device?

Keep it local when the recording is privacy-bound, when the connection is unreliable, or when you need a transcript in milliseconds inside an app. A push-to-talk note, a wake-phrase check or a short voice command all fit inside 30 seconds and one of the seven languages.

Cactus also reports that Whistle beats Whisper base on word error rate for several benchmarks (LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22, FLEURS average) while Whisper base is better on TED-LIUM, AMI and the MLS average. So the quality claim is dataset-specific. Test on your own audio before you commit.

When should you send it to Sume STT?

Send it to Sume when the clip is longer than one window, when the language is outside Whistle's list, or when the transcript feeds something else Sume does. Sume STT 1.0 is POST /v1/stt-1.0/transcribe; you submit an audio_url, an Idempotency-Key and an optional language_code or duration_seconds, then read the job result. words[] always comes back, and segmentation with mode sentence adds sentence segments.

For a clip that is still inside a video, video inspect with transcribe true runs the same STT at the same $0.01 per audio minute, and audio detach pulls the track out as wav or mp3 first. Jobs are asynchronous, so store the job id and poll its status rather than resubmitting a paid request.

How do you decide in practice?

Ask three questions in order.

  • Is the audio allowed to leave the device? If not, Whistle or another local model is the only answer.
  • Does it fit 30 seconds and one of the seven languages? If yes, local is the cheapest and fastest path.
  • Otherwise, upload the file, pay per minute, and take the word timings into captions or a timeline on Sume.

Sources

Related posts

More in Models

All Models posts

Written by Sume