Sume TTS 1.0 rejects model and model_id: use the router to pick
TTS 1.0 has no engine picker and returns 400 for model or model_id. The TTS router takes a required model from its catalog. Compare with ElevenLabs model tiers.

If you send model or model_id to Sume's POST /v1/tts-1.0/generate, you get a 400: TTS 1.0 has no engine picker and its public id is sume/tts-1.0. To choose an engine explicitly, call POST /v1/tts-router/generate with a required model from GET /v1/tts-router/models. The router uses the same transcript, voice and output rules and the same billing as TTS 1.0; the only difference is that the engine is your choice, and its documented engines are Cartesia Sonic ids.
Two routes, one difference
Both routes take the same transcript fields (up to 20,000 characters), the same voice selectors and the same output options. An unknown model id on the router fails with 400 model_not_found and the catalog URL, so the catalog is the source of truth, not a blog post. The OpenAPI enum currently names sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview; read the live list before hard-coding any of them.
| Question | TTS 1.0 | TTS router |
|---|---|---|
| Path | /v1/tts-1.0/generate | /v1/tts-router/generate |
| Engine field | Rejected with 400 | Required model |
| Where to list engines | Not applicable | GET /v1/tts-router/models |
| Billing | $0.0475 per 1,000 characters | Same per-character rate, list x 1.25 |
| Result | Hosted audio artifact via job URLs | Same |
How other vendors tier the choice
On ElevenLabs the model is part of the call and each tier has a job. The models overview says Eleven v4 is for professional content with 90+ languages and a 10,000-character limit, v4 Turbo is a real-time option at about 100 ms, Flash v2.5 is about 75 ms with 32 languages and a 40,000-character limit, and Multilingual v2 is the previous generation with 29 languages (read 2026-10-04). OpenAI separates tts-1, tts-1-hd and gpt-4o-mini-tts in its guide.
Sume's default hides the choice so most callers never pick wrong; the router exists for the callers who want to pin an engine for repeatability or to compare two.
When to pin an engine
Pin one when a series must sound the same across months, when you are auditioning engines for a language, or when you need a documented id in a log. Otherwise TTS 1.0 is the simpler call. Either way, read the result from the job as described in jobs and results, and keep the job.model value, which the router sets to the id you requested. The models overview lists the surfaces side by side.
Sources
Related posts
More in Developers
- Sume TTS emotion is a 64 character string: write a guide that fits
The emotion field in Sume TTS generation_config takes 1 to 64 characters. How to write a short, usable guide and test it against a neutral take.
- Sume TTS takes transcript or transcript_source, never both
A Sume TTS request accepts exactly one text input: literal transcript, or a transcript_source that points at a stored script. How to pick.
- Sume TTS volume 0.5 to 2.0: set the voiceover level before the mix
generation_config volume is a multiplier from 0.5 to 2.0 on Sume TTS. Use it to match narration loudness across jobs before you join or mix them.
- Check for a newer Sume CLI release in CI without auto-upgrading
sume update --check reports whether a newer GitHub Release exists and changes no files. Run it on a schedule, log it, and bump your pinned tag by pull request.
Written by Sume