Sume TTS 1.0 or TTS Router: which endpoint for a voiceover?
Both bill $0.0475 per 1,000 characters. TTS 1.0 always runs sonic-3.6; the router makes you name a Sonic id. Pick by whether you need to pin the engine.

Use POST /v1/tts-1.0/generate when you want a voiceover and do not care which engine version speaks; use POST /v1/tts-router/generate when you want the job to record the exact Sonic id you chose. Both take the same text, voice, language and output fields, and both bill the same character book, so the choice is about reproducibility, not price.
The two surfaces side by side
Sume's router doc describes the router as a public pass-through surface that is separate from Sume TTS 1.0. The difference is who picks the engine.
| Item | TTS 1.0 | TTS Router |
|---|---|---|
| Endpoint | POST /v1/tts-1.0/generate | POST /v1/tts-router/generate |
| Engine selection | Managed; voice or avatar only | Required model from the catalog |
job.model | Always sume/tts-1.0 | The id you requested, such as sonic-3.6 |
| Engine used | sonic-3.6 | The id you named |
| Price | $0.0475 per 1,000 characters | $0.0475 per 1,000 characters |
Unknown or extra model | model or model_id returns 400 | Unknown id returns 400 with catalog_url |
The five router ids
The router's first version is Cartesia Sonic only: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. Per the doc, sonic-latest is an alias that resolves to sonic-3.6, and sonic-preview is a beta channel whose output and availability can change without notice. It also rejects pro voice clones with voice_model_mismatch.
Eleven, OpenAI and other TTS families are not in the router. The doc says later vendors will be new catalog rows, not a new surface, and lists streaming TTS among the first ship's non-goals.
When to pin a router id
Use TTS 1.0 when none of that applies. It always uses sonic-3.6, the provider's current stable Sonic release, so you inherit upgrades only when Sume moves its own default.
- A campaign where every localized read must come from one engine version: name
sonic-3.6and the job record says so. - A blind comparison of
sonic-3.6againstsonic-3.5on your own script. Same price, same voice, one field changes. - A pipeline that should fail loudly when an id disappears: an unknown id returns 400 with the catalog URL instead of silently moving.
- Avoid
sonic-previewfor production reads. Its output can change without notice.
A router request
The voice selectors are the same as TTS 1.0: top-level avatar_id or avatar_handle, or voice.id. The router does not create a second voice namespace. Send exactly one of transcript (1 to 20,000 characters) or transcript_source.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: vo-router-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Welcome to the autumn sale.",
"voice": { "id": "YOUR_TTS_VOICE_UUID" },
"language": "en",
"timestamps": { "words": true }
}'Billing and polling are shared
The router uses the same admit, reserve and capture path as TTS 1.0, and the same GET /v1/jobs/:id/status and /result polling. Each catalog row is character-metered at Cartesia list micros times 1.25, rounded up to the cent, so a 1,350-character read is 7 cents on either surface and the minimum job is 1 cent.
Idempotency matters here. Send a distinct Idempotency-Key for each intended take, and reuse a key only to retry the same take, so a timeout never pays twice and a new wording never returns the old audio.
Because the price is identical, an A/B between engine ids costs only the extra takes. Two takes of a 1,350-character script are 14 cents.
A decision in three questions
If all three answers are no, TTS 1.0 is the simpler call: no model field to maintain, and the price is the same. You can move to the router later without changing voices, because the selectors are shared.
- Do you need the job record to name the engine? If yes, use the router.
- Do you need an engine other than Sonic? Neither surface offers one today, so look elsewhere or wait for a new catalog row.
- Do you need streaming audio? Neither surface offers it in this first version.
Sources
Related posts
More in Developers
- Sume TTS 400: send transcript or transcript_source, never both
A TTS request needs exactly one of transcript or transcript_source. Both, or neither, is an error. Live-commerce Formats need the source. A validator.
- Sume TTS has no SSML field: speed, emotion, pronunciation dictionary
Moving an SSML voice script to Sume TTS? The request takes plain transcript text, so use speed 0.6 to 1.5, volume, emotion and a pronunciation dictionary id.
- Sume /v1/videos says cancelled, /v1/jobs says canceled: guard it
The /v1/videos poll uses pending, in_progress and cancelled; /v1/jobs uses queued, processing and canceled. A small normalizer keeps your poller from hanging.
- Sume webhook and status poll race: never move a job row backwards
A late poll can say processing after the webhook already said completed. A rank-guarded SQLite update keeps a Sume job row from going backwards.
Written by Sume