Sume TTS Router: five Sonic ids, same price as TTS 1.0
Sume TTS Router needs a model: one of five Sonic ids. It uses the same $0.0475 per 1,000 characters as TTS 1.0. What the model field changes.

What does the Sume TTS Router add over TTS 1.0? One required field, model, chosen from five ids: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. The transcript, voice selector and output rules are the same, and both surfaces list $0.0475 per 1,000 characters in the Sume catalog.
These details come from the Sume OpenAPI spec. Use TTS 1.0 when you do not care which engine runs, and the router when you need to pin one.
Side by side
TTS 1.0 has no engine picker: sending model or model_id to it returns a 400 and tells you to use the router. The router does the reverse and requires model. Unknown ids fail with 400 model_not_found and the catalog URL, and job.model echoes the id you asked for.
| Item | TTS 1.0 | TTS Router |
|---|---|---|
| Endpoint | POST /v1/tts-1.0/generate | POST /v1/tts-router/generate |
| model field | Rejected with 400 | Required, one of five ids |
| Catalog price | $0.0475 per 1,000 characters | $0.0475 per 1,000 characters |
| Transcript limit | 20,000 characters | Same 20,000 today; max_characters is listed per model |
| Discover models | Not applicable | GET /v1/tts-router/models |
Reading the router catalog
GET /v1/tts-router/models returns each model with its capabilities, including a max_characters integer, a per-character list price in USD micros, a billable margin field and a constraints list. Today every router id lists the same 20,000-character max_characters as TTS 1.0, but read the catalog rather than hard-coding limits, since the field is per model.
The catalog says the router uses the same character book as TTS 1.0, so a switch between the two does not change what a given script costs.
Which one to use
Pick on whether the engine is part of your contract.
- Use TTS 1.0 for most narration. It chooses the engine for you, so a catalog change does not break your code.
- Use the router when a customer or a test needs a specific engine, or when you want to compare two ids on the same script.
- Use
sonic-latestandsonic-previewknowing they move: the catalog describessonic-latestas an alias forsonic-3.6andsonic-previewas a provider beta channel whose output and availability can change without notice, and not compatible with pro voice clones (those requests fail withvoice_model_mismatch). Pin a numbered id for reproducible audio.
What stays the same
Voice selection, language, output_format, timestamps and sentence segmentation work the same on both. A voice comes from an avatar reference or a voice id in UUID or voi_ shape. Results arrive as a job, read by polling or webhook as described in Sume jobs and results; sync mode waits at most 30 seconds.
Check the output container before you mux: WAV vs MP3 before an avatar mux explains why WAV is the safer default for that step.
A pinning test
To compare two engines fairly, send the same transcript, the same voice and the same output format to the router twice, changing only model. Keep the job.model value from each result with the audio file name. Because the catalog prices the engines the same way, the comparison costs the same per character whichever pair you test, so the choice rests on how the audio sounds and how it fits your edit. Listen on the playback device your audience uses, not only on studio headphones.
Sources
Related posts
More in Models
- Sume image models at a $0.02 list price: Grok, Qwen, Imagen Fast
Three Sume image models list at $0.02 per image: Grok Imagine, Qwen Image and Imagen 4 Fast. How they differ on edits, ratios and image count per call.
- Veda sparse attention for MiniMax H3: 6.8x attention, 3.1x clip
Veda's sparse attention keeps 10 percent of attention work for MiniMax H3. Why 6.8x on attention becomes 3.1x per clip, and what it needs to run.
- Gemini 3.8 Flash is stable: keep model ids in config
When a vendor ships a new Flash model, a hard-coded id ages. Read Sume ids from GET /v1/catalog and treat 404 model_not_found as a signal, not a retry.
- How long can a Veo 3.1 video get? 148 seconds at 720p
Veo 3.1 extends clips 7 seconds at a time, up to 148 seconds at 720p. Sume has no extend task, so here are the longer single-clip models and how to join clips.
Written by Sume