Phonon-2 (164 MB, English) or hosted Sume STT for a 10-minute file?
Phonon-2 is a free 164 MB English-only on-device model. Sume STT is hosted, multilingual, about 1 cent a minute. When each fits, with sourced numbers.

Pick Phonon-2 if your audio is English, the user's Mac or iPhone can run a 164 MB model, and you can live with 35-second utterances. Pick hosted Sume STT if you need other languages, a 10-minute file in one request, or a transcript without shipping a model. Fermion Research released Phonon-2 on September 29, and its page lists a 5.21% average word error rate over seven English test sets.
The trade is not accuracy against price. It is where the work runs. A local model costs you no per-minute fee but you carry the download, the update path and the per-device variance. A hosted job costs about one cent a minute and moves audio to a server. Neither is free of work; you choose which work you own.
The numbers side by side
Phonon-2 figures come from the vendor's own page. Sume figures come from the API schema and the pricing tables in the repository: the provider list price is $0.008 a minute and Sume applies its 1.25 standard margin, which gives about $0.01 a minute.
| Item | Phonon-2 | Sume STT |
|---|---|---|
| Where it runs | On device (Core ML) | Hosted job |
| Download | 164 MB | None |
| Languages | English only | Many; omit language_code to auto-detect |
| Longest input | Utterances up to 35 s (Core ML) | 10 minutes per request |
| Streaming | No, non-streaming | No, async job |
| Word timestamps | Per-word | Word timings always returned |
| Reported accuracy | 5.21% average WER, seven English sets | Not published here; measure your own clips |
| Cost | Free to run; CC-BY-4.0 licence | About $0.01 per minute |
Run the hosted path in four steps
- Upload the file so it has a public HTTPS address; a Sume media URL is preferred.
- POST it to /v1/stt-1.0/transcribe with an Idempotency-Key and Bearer key.
- Send duration_seconds (1 to 600) so the reservation matches the clip; leaving it out reserves one minute.
- Read job.id, status_url and result_url from the response, then poll /v1/jobs/:id/status and fetch /v1/jobs/:id/result.
import os, requests
r = requests.post(
"https://api.sume.com/v1/stt-1.0/transcribe",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "stt-demo-001",
},
json={
"audio_url": "https://media.sume.com/example/clip.mp3",
"duration_seconds": 540,
},
timeout=30,
)
r.raise_for_status()
print(r.json())Where the hosted job helps
A 10-minute interview is 600 seconds, far past the 35-second utterance size that the Core ML build handles, so an app that wants a whole file in one go would have to cut the audio into chunks and stitch the words back together. The hosted job takes the file whole and returns word timings you can use for captions or search. The cost is easy to predict: a 10-minute file is about 10 minutes of billing, which is roughly 10 cents at the rate above.
The local model has a real edge on ongoing volume. If you transcribe thousands of hours of English every month on devices you already own, the marginal cost is close to nothing. Do the arithmetic for your own volume before you choose; a few hundred minutes a month will not repay the engineering time of shipping and updating a model.
What Sume does not do
Sume STT does not run on your device, does not stream, does not label speakers (diarization and audio-event tagging are fixed off on the server), and takes at most 10 minutes of audio per request. The docs also note that newer generators can appear on the dev API before the production OpenAPI snapshot, so check your own account before you plan around it.
How to decide
Run both on 20 of your own clips and count the words that matter, such as names and numbers. Our guide to measuring word error rate on your own clips has a script. If the local model clears your bar, ship it. If your audio is not English or runs past 35 seconds, the hosted job is the simpler path. For the privacy side of the same choice, see where dictation audio goes.
Sources
Related posts
More in Comparisons
- Photoroom Plus $0.10 vs Basic $0.02: what Sume covers
Photoroom Basic removes backgrounds at $0.02 an image; Plus adds AI backgrounds, shadows and try-on at $0.10. Which of those Sume covers, and which it does not.
- Pika Starter 900 credits vs Creator 3,150: same credit price
Pika Creator costs 3.5 times Starter and gives 3.5 times the credits. The extra $25 is the commercial license, not a volume discount.
- Placid 1 credit and Bannerbear 2 per video second vs Sume
Placid charges one credit per second of video and Bannerbear two per second of animation. Sume Timeline charges $0.10 per output minute. Convert them.
- Poster text test: Nano Banana Pro at 19 cents vs GPT Image 2.5 at 7
A 45-cent test plan for poster text: run one prompt on Nano Banana Pro (19 cents) and GPT Image 2.5 high (7 cents) on Sume, then judge spelling and layout.
Written by Sume