Gradium's 1,421 voices: how you pick a voice on Sume TTS
Gradium listed 1,421 TTS voices on Oct 7. Sume does not carry Gradium; here is how voice selection works on Sume TTS with avatar handles or voice ids.

Sume does not offer Gradium voices. Gradium's own page says its catalog grew to 1,421 voices on 2026-10-07, while Sume's text-to-speech route runs on the Sonic engines in the TTS Router catalog and picks the voice from an avatar or a voice id, not from a browsable catalog of that size.
This page compares what Gradium says it shipped with how a voice is chosen on Sume, using Gradium's announcement (read 2026-10-10) and Sume's published API schema from the API reference. If you came here wanting to switch TTS vendors, the useful question is whether your workflow needs a large browsable voice list or a stable, repeatable voice per brand. Sume is built for the second.
What did Gradium announce?
Gradium's post, titled "1,000+ new voices, our most natural set yet", says the catalog went from roughly 387 voices to 1,421. It lists five languages (English, French, Spanish, German, Portuguese) and 28 accents, and groups voices by use: customer service, narration, ads and social, and character voices. It also quotes sub-50 ms latency for the text-to-speech model. The page gives no price and no licensing terms, so this post makes no cost claim about Gradium.
| Item | What the page says |
|---|---|
| Total voices | 1,421 (up from about 387) |
| Languages | 5: English, French, Spanish, German, Portuguese |
| Accents | 28 |
| Voices tagged for narration | 221 |
| Voices tagged for ads and social | 228 |
| Voices tagged for customer service | 256 |
| Voices tagged as characters | 329 |
| Latency claim | Sub-50 ms |
How does Sume pick a voice?
Sume TTS 1.0 needs a transcript and one voice selector. The documented selectors are the top-level avatar_id or avatar_handle (the voice of one of your workspace avatars, usable when the avatar's voice.status is ready) or voice.id, a Sume TTS voice UUID or a Voices library id that starts with voi_.
A voice name from another vendor is not accepted. The API rejects any other id shape with 400 invalid_voice_id before a job is queued and before credits are reserved, so a wrong voice name costs nothing. If you hold only the API contract, list your avatars at GET /v1/avatar-1.0/avatars and use one whose voice is ready.
avatar_idoravatar_handle: resolves the avatar's voice at submit time.voice.id: a TTS voice UUID, orvoi_plus 32 hex characters from your Voices library.- If you send an avatar and
voice.idtogether, they must resolve to the same voice, or the request fails with 400.
Which engine says what the voice sounds like?
The engine is separate from the voice. TTS 1.0 has no engine picker and rejects model. To pin an engine, call POST /v1/tts-router/generate with a required model, one of sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview. sonic-latest is documented as an alias for sonic-3.6, never sonic-preview. All of them use the same character price, $0.0475 per 1,000 characters.
| Field | Where | Role |
|---|---|---|
avatar_id / avatar_handle | TTS 1.0 and TTS Router | Pick the voice from a ready avatar |
voice.id | TTS 1.0 and TTS Router | Pick a voice by id |
model | TTS Router only | Pick the Sonic engine (required there, rejected on TTS 1.0) |
language | Both | Language the voice speaks the transcript in |
When is a bigger catalog the wrong thing to chase?
A catalog of 1,421 voices helps when a person picks a voice by ear once. Automated pipelines do the opposite: they need the same voice on every episode. On Sume the avatar or voice id is that pin, and the job result stores the audio under Sume media URLs you can reuse. Compare that with browsing: the more voices you can pick from, the more you have to document which one your brand uses.
If a named Gradium voice is a hard requirement, Sume is the wrong tool for that step; use Gradium directly, import the finished audio, and use Sume for the parts it does list, such as captions, timeline assembly and lip-sync clips. Imported audio must be public HTTPS or already on the Sume media host depending on the route, so check the route's audio field before you plan the handoff.
Sources
Related posts
More in Models
- Gradium TTS covers 5 languages: what Sume's language field does
Gradium lists English, French, Spanish, German and Portuguese. Sume TTS takes a language field and a mismatch guard; here is how a non-English line behaves.
- Gradium's sub-50 ms TTS vs Sume's async TTS jobs: which fits
Gradium claims sub-50 ms latency. Sume TTS is a non-streaming job with a 30 second wait window; it suits batch narration and video, not live agent replies.
- Grok Imagine on Sume lists 9:19.5 and 9:20 but not 4:5
Sume's Grok Imagine row lists tall phone ratios 9:19.5 and 9:20 and wide 20:9 and 19.5:9, but not 4:5, 5:4 or 21:9, and caps n at 1. Here is what to do.
- Grok Imagine multi-image edit: xAI says 5 sources, Sume lists 10
xAI's Imagine docs cap multi-image edits at 5 source images and 10 outputs per request. Sume's Grok Imagine row lists 10 references and one output. Port safely.
Written by Sume