Whistle's 30-second window: passes for 5 minutes vs Sume STT
A 5-minute recording is ten 30-second Whistle passes on the device, or one Sume STT job at $0.05. How to plan chunking and when to skip it.

Whistle takes at most 30 seconds of audio per pass, so a 5-minute recording needs at least ten passes and a 10-minute recording at least twenty. Sume STT takes the whole file as one job: 5 minutes is $0.05 and 10 minutes is $0.10 at the $0.01 per audio minute rate. The extra work in the local route is your chunking code, not the model.
The 30-second figure and the 16 kHz mono input are from the Cactus Compute launch post for Whistle (read 2026-10-11). Cactus presents Whistle as an on-device model of 16.9 MB; the part of the post I read does not describe stitching long audio, so plan on cutting it yourself.
How many Whistle passes does a recording need?
Divide the length in seconds by 30 and round up. That is the minimum, before you add any overlap. The table shows the arithmetic next to the Sume job price; the passes are computed from the 30-second limit, and the Sume figures are the published rate times the length.
| Recording | Minimum Whistle passes | Sume STT 1.0 jobs | Sume STT cost at $0.01 per minute |
|---|---|---|---|
| 30 seconds | 1 | 1 | $0.01 (the catalog minimum is 1 cent) |
| 1 minute | 2 | 1 | $0.01 |
| 5 minutes | 10 | 1 | $0.05 |
| 10 minutes | 20 | 1 | $0.10 |
What goes wrong when you chunk speech?
A hard cut every 30 seconds will often land in the middle of a word. Words at the edge of a window can be dropped or split, and the second pass has no memory of the first. Cut on silence instead, or overlap the windows by a second or two and drop duplicated words when you merge.
Word timings add one more step. Each pass reports times relative to its own window, so add the window's start offset before you merge. Forgetting the offset gives captions that reset to zero every 30 seconds. Sume's own transcript returns times for the whole file, so there is nothing to offset.
What does Sume STT need from you?
Submit POST /v1/stt-1.0/transcribe with a public HTTPS audio_url and an Idempotency-Key. Pass duration_seconds if you know it, from 1 to 600; without it Sume reserves for one minute, and the catalog states a maximum of 10 minutes per request. The job is asynchronous, so keep the job id and poll its status; do not submit the paid request again because your own process timed out.
If the audio is the soundtrack of a video, audio detach cuts it out as wav or mp3 first, and video inspect can run the transcript in the same call. A silent clip is refused with inspect_source_has_no_audio, so check the probe first.
Which route should a 5-minute recording take?
If the audio must stay on the device, chunk for Whistle and test your merge on a recording with speech across the boundary. If it can go to the cloud, one Sume job is simpler and its price is easy to predict. Long meetings beyond 10 minutes need splitting on either route.
Sources
Related posts
More in Models
- Whistle keyword biasing vs Sume STT: fixing product names in captions
Whistle can bias decoding toward keywords. Sume STT takes no vocabulary list, so for captions you pass script_text instead. How the two fixes compare.
- MiMo V2.6 Pro and Flash in Sume's agent picker: what the catalog says
Sume's model catalog has rows for Xiaomi MiMo V2.6 Pro and Flash behind the OpenRouter switch. What the repo records, and what the API cannot pick.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
Written by Sume