TTS latency vs quality: Sume has async jobs, no streaming tier

Sume TTS is an async file job on Sonic models at one price. There is no streaming or latency tier, so you trade by script size and wait mode, not model.

4 min readSume
All posts

Sume TTS does not offer a latency-versus-quality dial. Every model on the router is priced the same, the router docs list streaming TTS as out of scope, and jobs are asynchronous files. If you want first audio in a few hundred milliseconds while text is still arriving, that is a live voice stack, not Sume TTS. If you want a finished narration file, the trade-offs below are the ones you actually have.

What you can and cannot tune

Quality-oriented choices on Sume are model id, voice, language, speed, volume and emotion. None of them is a faster, cheaper tier. Speed, volume and emotion live in generation_config, and a separate speed value accepts slow, normal or fast. The model choice is on the router only. Because price per character is identical across Sonic ids, a newer model does not cost more and an older one does not cost less.

Sume rates and limits, read 2026-10-06 from the repo catalog and schemas
LeverAvailable on Sume?Effect
Streaming audio outNo, listed as a non-goal for the routerYou get a finished file
Faster or cheaper model tierNo, one per-character ratePick by listening
Model idYes, on /v1/tts-router/generatesonic-3.6, 3.5, 3, latest, preview
Wait for the result inlineYes, sync mode up to 30 secondsShort lines only
Webhook on completionYesLong scripts

Where the real latency comes from

Mostly script length. A 100-character line and a 20,000-character script are the same API call with very different waits. For long text, split it into sentence-bounded chunks and submit them in parallel if your plan's concurrency allows, or use sentence segmentation to get boundaries inside one file. Measure your own numbers: time from submit to a finished file, on your own script, at the hour you ship. This post does not quote a figure because the docs do not publish one.

A practical decision rule

Ask what the listener does while the audio is being made. If a person is waiting on the line, you need streaming and Sume is not it. If a render queue is waiting, you need a finished file, and the question becomes cost and consistency rather than milliseconds. Sume's price per character is the same for every Sonic id, so consistency is the lever you actually control: pin the model, pin the voice, and keep the settings in the record.

When to look elsewhere

Phone agents, live translation and anything interactive need streaming in both directions. Use a streaming vendor for the call itself, and use Sume for the assets around it: an intro read, hold messages, post-call transcripts and the finished marketing video. The earlier post on when Sume TTS is the wrong tool lists those cases.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume