Phonon-2 (164 MB, English) or hosted Sume STT for a 10-minute file?

Phonon-2 is a free 164 MB English-only on-device model. Sume STT is hosted, multilingual, about 1 cent a minute. When each fits, with sourced numbers.

4 min readSume
All posts

Pick Phonon-2 if your audio is English, the user's Mac or iPhone can run a 164 MB model, and you can live with 35-second utterances. Pick hosted Sume STT if you need other languages, a 10-minute file in one request, or a transcript without shipping a model. Fermion Research released Phonon-2 on September 29, and its page lists a 5.21% average word error rate over seven English test sets.

The trade is not accuracy against price. It is where the work runs. A local model costs you no per-minute fee but you carry the download, the update path and the per-device variance. A hosted job costs about one cent a minute and moves audio to a server. Neither is free of work; you choose which work you own.

The numbers side by side

Phonon-2 figures come from the vendor's own page. Sume figures come from the API schema and the pricing tables in the repository: the provider list price is $0.008 a minute and Sume applies its 1.25 standard margin, which gives about $0.01 a minute.

Phonon-2 versus Sume STT (read 2026-10-08)
ItemPhonon-2Sume STT
Where it runsOn device (Core ML)Hosted job
Download164 MBNone
LanguagesEnglish onlyMany; omit language_code to auto-detect
Longest inputUtterances up to 35 s (Core ML)10 minutes per request
StreamingNo, non-streamingNo, async job
Word timestampsPer-wordWord timings always returned
Reported accuracy5.21% average WER, seven English setsNot published here; measure your own clips
CostFree to run; CC-BY-4.0 licenceAbout $0.01 per minute

Run the hosted path in four steps

  • Upload the file so it has a public HTTPS address; a Sume media URL is preferred.
  • POST it to /v1/stt-1.0/transcribe with an Idempotency-Key and Bearer key.
  • Send duration_seconds (1 to 600) so the reservation matches the clip; leaving it out reserves one minute.
  • Read job.id, status_url and result_url from the response, then poll /v1/jobs/:id/status and fetch /v1/jobs/:id/result.
import os, requests

r = requests.post(
    "https://api.sume.com/v1/stt-1.0/transcribe",
    headers={
        "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Idempotency-Key": "stt-demo-001",
    },
    json={
        "audio_url": "https://media.sume.com/example/clip.mp3",
        "duration_seconds": 540,
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

Where the hosted job helps

A 10-minute interview is 600 seconds, far past the 35-second utterance size that the Core ML build handles, so an app that wants a whole file in one go would have to cut the audio into chunks and stitch the words back together. The hosted job takes the file whole and returns word timings you can use for captions or search. The cost is easy to predict: a 10-minute file is about 10 minutes of billing, which is roughly 10 cents at the rate above.

The local model has a real edge on ongoing volume. If you transcribe thousands of hours of English every month on devices you already own, the marginal cost is close to nothing. Do the arithmetic for your own volume before you choose; a few hundred minutes a month will not repay the engineering time of shipping and updating a model.

What Sume does not do

Sume STT does not run on your device, does not stream, does not label speakers (diarization and audio-event tagging are fixed off on the server), and takes at most 10 minutes of audio per request. The docs also note that newer generators can appear on the dev API before the production OpenAPI snapshot, so check your own account before you plan around it.

How to decide

Run both on 20 of your own clips and count the words that matter, such as names and numbers. Our guide to measuring word error rate on your own clips has a script. If the local model clears your bar, ship it. If your audio is not English or runs past 35 seconds, the hosted job is the simpler path. For the privacy side of the same choice, see where dictation audio goes.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume