TTS latency vs quality: Sume has async jobs, no streaming tier
Sume TTS is an async file job on Sonic models at one price. There is no streaming or latency tier, so you trade by script size and wait mode, not model.

Sume TTS does not offer a latency-versus-quality dial. Every model on the router is priced the same, the router docs list streaming TTS as out of scope, and jobs are asynchronous files. If you want first audio in a few hundred milliseconds while text is still arriving, that is a live voice stack, not Sume TTS. If you want a finished narration file, the trade-offs below are the ones you actually have.
What you can and cannot tune
Quality-oriented choices on Sume are model id, voice, language, speed, volume and emotion. None of them is a faster, cheaper tier. Speed, volume and emotion live in generation_config, and a separate speed value accepts slow, normal or fast. The model choice is on the router only. Because price per character is identical across Sonic ids, a newer model does not cost more and an older one does not cost less.
| Lever | Available on Sume? | Effect |
|---|---|---|
| Streaming audio out | No, listed as a non-goal for the router | You get a finished file |
| Faster or cheaper model tier | No, one per-character rate | Pick by listening |
| Model id | Yes, on /v1/tts-router/generate | sonic-3.6, 3.5, 3, latest, preview |
| Wait for the result inline | Yes, sync mode up to 30 seconds | Short lines only |
| Webhook on completion | Yes | Long scripts |
Where the real latency comes from
Mostly script length. A 100-character line and a 20,000-character script are the same API call with very different waits. For long text, split it into sentence-bounded chunks and submit them in parallel if your plan's concurrency allows, or use sentence segmentation to get boundaries inside one file. Measure your own numbers: time from submit to a finished file, on your own script, at the hour you ship. This post does not quote a figure because the docs do not publish one.
A practical decision rule
Ask what the listener does while the audio is being made. If a person is waiting on the line, you need streaming and Sume is not it. If a render queue is waiting, you need a finished file, and the question becomes cost and consistency rather than milliseconds. Sume's price per character is the same for every Sonic id, so consistency is the lever you actually control: pin the model, pin the voice, and keep the settings in the record.
When to look elsewhere
Phone agents, live translation and anything interactive need streaming in both directions. Use a streaming vendor for the call itself, and use Sume for the assets around it: an intro read, hold messages, post-call transcripts and the finished marketing video. The earlier post on when Sume TTS is the wrong tool lists those cases.
Sources
Related posts
More in Comparisons
- Unreal Speech TTS plans in dollars per million characters vs Sume
Unreal Speech's plans work out from $16.33 to $8.00 per million characters. Sume lists $47.50 with no plan. Break-even by monthly volume.
- Vadoo cost per video: $0.20 to $0.95 a plan vs Sume per clip
Vadoo plans work out to $0.20 to $0.95 per video included, billed monthly or yearly. Sume prices each AI video clip by model and length instead.
- Veo 3.1 Lite 1080p at 8 cents a second vs Sume 1080p models
Google lists Veo 3.1 Lite 1080p at $0.08 a second. Here is the same 10-second 1080p clip priced on Veo tiers and on Sume models, with the gaps stated plainly.
- Which AI video models charge the same for 720p and 1080p?
Google lists Veo 3.1 Standard at $0.40 a second for both 720p and 1080p, and Sume's Kling Video v3 Pro has one rate for both. Other models charge more at 1080p.
Written by Sume