Does dictation audio leave your device? Fireflies, Phonon-2, Sume
Fireflies Talk keeps finished dictations local but sends audio to its STT provider. Phonon-2 runs on device. Sume STT is hosted. A sourced privacy comparison.

It depends on the product. Fireflies Talk, launched on October 1, stores completed dictations locally but sends the audio to its speech-to-text provider, according to Fossbytes. Phonon-2 is a 164 MB model that runs on Apple hardware, so audio need not leave the device. Sume STT is a hosted job, so audio is fetched from a URL you supply.
So the useful question is not private or not private. It is: who receives the raw audio, and who keeps the text. Write that down for each tool before you pick one.
Three designs, three answers
The table lists what each source states. Where a source is silent I say so rather than guess.
- Fireflies lists SOC 2 Type II, GDPR and HIPAA compliance, per Fossbytes.
- Phonon-2 is English only and non-streaming.
- Sume STT takes up to 10 minutes of audio and returns word timings.
| Tool | Audio | Text |
|---|---|---|
| Fireflies Talk | Sent to its STT provider | Completed dictations stored locally |
| Phonon-2 | On device, per the vendor page | Stays in your app |
| Sume STT | Fetched from your audio_url by a hosted job | Returned in the job result |
How to run a privacy check in four steps
- List every hop the audio takes, from microphone to final text.
- Ask each vendor for its retention terms in writing; this post does not state them.
- Decide whether the audio may sit at a public HTTPS address, which Sume STT requires.
- Delete the source file from storage after the job result is saved.
- Write the answers into a one-page note that your team can review each quarter.
Questions to put to every vendor
Ask where the audio is processed, how long it is kept, who can read it, and whether it is used to train a model. Ask for the answers in writing, with the page or contract that states them. For on-device models, ask what the app sends back, if anything, such as crash logs or usage events.
Then compare the answers with what your own users were told. A mismatch between a privacy notice and the real data path is the usual cause of incidents, not the model itself. If a recording holds sensitive material, keep a record of the decision and the date, so you can show why a given tool was chosen.
What Sume does not do
Sume STT needs a public HTTPS audio_url, and the docs prefer a Sume media URL. It has no device-side mode and no live microphone input. This post does not describe Sume retention rules, because the pages I read do not state them. Ask for them before you send sensitive audio.
Pick by risk
For private dictation of short English notes, on-device is simplest. For recorded interviews in many languages, a hosted job is practical. See the cost of transcribing 1,000 hours for the hosted side of the bill, and Fireflies Talk versus an API transcript of a recorded file for use cases.
Sources
Related posts
More in Comparisons
- Duck a Suno or ElevenLabs track under a voice-over in Sume Timeline
Sume Timeline mixes a soundtrack bed under the voice spine with duck_db from 0 to 20, loop and a fade of up to 10 seconds, at $0.10 per output minute.
- ElevenLabs Free 10,000 credits is 11 minutes of music, no license
ElevenLabs Free gives 10,000 credits a month, about 11 minutes of music, without a commercial license. Sketch there, then ship a track via Sume at $0.125.
- ElevenLabs Creator 121,000 shared credits vs Sume per-job music
ElevenLabs Creator ($22) gives 121,000 credits shared across products: 1 per speech character, 900 per music minute. Sume music is $0.125 per track.
- Scribe v2 Realtime $0.39/h vs batch $0.22/h: when does Sume STT fit?
ElevenLabs lists Scribe v2 at $0.22 an hour and Scribe v2 Realtime at $0.39. How 1,000 hours compares with Sume STT at about $0.60 an hour.
Written by Sume