Is there an ElevenLabs or OpenAI model in Sume's TTS Router? No
Sume's TTS Router v1 lists Cartesia Sonic models only. Other vendors would arrive as new catalog rows. How to check the catalog and what the 400 looks like.

No. The TTS Router on Sume currently lists Cartesia Sonic only: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. The router docs state that v1 does not include Eleven, OpenAI or other TTS families, and that later vendors would be new catalog rows rather than a new surface.
What the router is
The router is a pass-through surface: you send a required model id from the catalog and Sume bills the listed per-character rate with its margin. TTS 1.0 is the managed product with no model field; if you send model or model_id there, you get a 400. On the router, an unknown model returns a 400 that includes a catalog_url pointing at /v1/tts-router/models.
| Question | Answer |
|---|---|
| Which vendors? | Cartesia Sonic only in v1 |
| Streaming TTS? | Out of scope for the first ship |
| Routing preset? | Out of scope |
| A model picker in the Agents or Studio UI? | Out of scope |
| How do new models appear? | As new rows in the catalog endpoint |
How to check before you build
Call GET /v1/tts-router/models and read the ids. Do not hard-code a list from a blog post, including this one. If your workflow depends on a specific vendor's voice, the catalog is the only reliable answer on a given day. A cheap way to stay informed is to poll the catalog and alert when an id appears.
Why a narrow catalog can be fine
A short catalog means fewer decisions. Every id has the same per-character price, so you can compare them on quality alone, and TTS 1.0 removes even that choice by always using sonic-3.6. The cost of a narrow catalog is that you cannot pick a different vendor's voice quality, language coverage or emotive range, so decide whether that matters for your project before you build on it.
What to do if you need another vendor
Use that vendor's own API for the voice, then bring the audio into Sume for the parts it does cover: captions, trimming and timeline assembly. Keep the vendor's price and terms in your own records, since Sume's rate card applies only to the models in its catalog.
Sources
Related posts
More in Comparisons
- Kling 4.0 or Seedance 2.5 for a 30-second AI video today?
Both advertise 30 seconds. Only one has a callable id on Sume today. A spec-by-spec read of the vendor pages and what you can actually request.
- Lemonfox TTS at $2.50 per million characters vs Sume: the honest gap
Lemonfox's page works out to about $2.50 per million characters. Sume lists $47.50. Cost for 900, 22,500 and 1M characters, and what the gap buys you.
- Lip-sync a video you have, or animate a photo: which API?
Re-syncing a mouth in footage and making a still talk are different jobs. A guide using Sync.so docs and Sume lip-sync, avatar and face-swap routes.
- LTX-2.5 or a lip-sync API for a talking head: what Sume offers
Sume does not list LTX-2.5. For a talking head that must say your exact words, the route that ships is TTS plus H3 Max lip sync, or Avatar 1.0 talking video.
Written by Sume