TTS 1.0 rejects a model field: use the TTS router to pick sonic-3.6
Sume TTS 1.0 does not accept model. To choose an engine such as sonic-3.6 use /v1/tts-router/generate. Price is the same $0.0475 per 1,000 characters.

TTS 1.0 on Sume rejects a model field, so to choose an engine such as sonic-3.6 you call the TTS Router at /v1/tts-router/generate instead. Both are billed at $0.0475 per 1,000 characters, so a 1,000-character script is $0.0475 either way.
Which route takes which fields
The router request schema is the TTS 1.0 body plus an optional model.
| Route | model field | Example |
|---|---|---|
| POST /v1/tts-1.0/generate | rejected | default route |
| POST /v1/tts-router/generate | optional | sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview |
Router request
Everything else is the same body you already use. Fetch the catalog before pinning an id.
curl https://api.sume.com/v1/tts-router/models \
-H "Authorization: Bearer $SUME_API_KEY"
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: router-demo-01" \
-d '{"model": "sonic-3.6", "transcript": "Hello from the router.",
"avatar_handle": "@narrator", "language": "en"}'Gotchas
sonic-latest and sonic-preview are moving targets; pin sonic-3.6 if you need repeatable output. Unknown ids fail with a model error that points at the catalog. Check the catalog response for per-model constraints before you assume a feature works on every engine.
Sources
Related posts
More in Developers
- Sume TTS pace test: 1,000 characters a minute decides the cap
Sume TTS stops at 20,000 characters or 1,200 seconds. The break-even is 1,000 characters a minute; measure your voice before a long read. Max $0.95 a job.
- Sume TTS speed 0.6 to 1.5: a 14-minute script and the 1,200 s cap
generation_config.speed runs 0.6 to 1.5. A 14-minute read slowed to 0.6 would need about 23 minutes and hit the 1,200-second cap; the price stays per character.
- Sume TTS: 20,000 characters or 1,200 seconds, which fails first?
At 15 characters a second, 20,000 characters is 1,333 seconds, so the 1,200-second audio cap trips first. Split near 15,000 characters per job.
- Sume TTS default is mp3: sentence slices need wav and emit_audio
Sume TTS defaults to mp3 at 44.1 kHz and 128 kbps. Per-sentence audio slices need wav (pcm_s16le) plus segmentation emit_audio, as in the Python request below.
Written by Sume