Sume TTS Router: five Sonic ids, same price as TTS 1.0

Sume TTS Router needs a model: one of five Sonic ids. It uses the same $0.0475 per 1,000 characters as TTS 1.0. What the model field changes.

5 min readSume
All posts

What does the Sume TTS Router add over TTS 1.0? One required field, model, chosen from five ids: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. The transcript, voice selector and output rules are the same, and both surfaces list $0.0475 per 1,000 characters in the Sume catalog.

These details come from the Sume OpenAPI spec. Use TTS 1.0 when you do not care which engine runs, and the router when you need to pin one.

Side by side

TTS 1.0 has no engine picker: sending model or model_id to it returns a 400 and tells you to use the router. The router does the reverse and requires model. Unknown ids fail with 400 model_not_found and the catalog URL, and job.model echoes the id you asked for.

TTS 1.0 and TTS Router compared from the Sume spec and catalog (read 2026-10-03)
ItemTTS 1.0TTS Router
EndpointPOST /v1/tts-1.0/generatePOST /v1/tts-router/generate
model fieldRejected with 400Required, one of five ids
Catalog price$0.0475 per 1,000 characters$0.0475 per 1,000 characters
Transcript limit20,000 charactersSame 20,000 today; max_characters is listed per model
Discover modelsNot applicableGET /v1/tts-router/models

Reading the router catalog

GET /v1/tts-router/models returns each model with its capabilities, including a max_characters integer, a per-character list price in USD micros, a billable margin field and a constraints list. Today every router id lists the same 20,000-character max_characters as TTS 1.0, but read the catalog rather than hard-coding limits, since the field is per model.

The catalog says the router uses the same character book as TTS 1.0, so a switch between the two does not change what a given script costs.

Which one to use

Pick on whether the engine is part of your contract.

  • Use TTS 1.0 for most narration. It chooses the engine for you, so a catalog change does not break your code.
  • Use the router when a customer or a test needs a specific engine, or when you want to compare two ids on the same script.
  • Use sonic-latest and sonic-preview knowing they move: the catalog describes sonic-latest as an alias for sonic-3.6 and sonic-preview as a provider beta channel whose output and availability can change without notice, and not compatible with pro voice clones (those requests fail with voice_model_mismatch). Pin a numbered id for reproducible audio.

What stays the same

Voice selection, language, output_format, timestamps and sentence segmentation work the same on both. A voice comes from an avatar reference or a voice id in UUID or voi_ shape. Results arrive as a job, read by polling or webhook as described in Sume jobs and results; sync mode waits at most 30 seconds.

Check the output container before you mux: WAV vs MP3 before an avatar mux explains why WAV is the safer default for that step.

A pinning test

To compare two engines fairly, send the same transcript, the same voice and the same output format to the router twice, changing only model. Keep the job.model value from each result with the audio file name. Because the catalog prices the engines the same way, the comparison costs the same per character whichever pair you test, so the choice rests on how the audio sounds and how it fits your edit. Listen on the playback device your audience uses, not only on studio headphones.

Sources

Related posts

More in Models

All Models posts

Written by Sume