Gradium's 1,421 voices: how you pick a voice on Sume TTS

Gradium listed 1,421 TTS voices on Oct 7. Sume does not carry Gradium; here is how voice selection works on Sume TTS with avatar handles or voice ids.

5 min readSume
All posts

Sume does not offer Gradium voices. Gradium's own page says its catalog grew to 1,421 voices on 2026-10-07, while Sume's text-to-speech route runs on the Sonic engines in the TTS Router catalog and picks the voice from an avatar or a voice id, not from a browsable catalog of that size.

This page compares what Gradium says it shipped with how a voice is chosen on Sume, using Gradium's announcement (read 2026-10-10) and Sume's published API schema from the API reference. If you came here wanting to switch TTS vendors, the useful question is whether your workflow needs a large browsable voice list or a stable, repeatable voice per brand. Sume is built for the second.

What did Gradium announce?

Gradium's post, titled "1,000+ new voices, our most natural set yet", says the catalog went from roughly 387 voices to 1,421. It lists five languages (English, French, Spanish, German, Portuguese) and 28 accents, and groups voices by use: customer service, narration, ads and social, and character voices. It also quotes sub-50 ms latency for the text-to-speech model. The page gives no price and no licensing terms, so this post makes no cost claim about Gradium.

Facts from Gradium's page (read 2026-10-10)
ItemWhat the page says
Total voices1,421 (up from about 387)
Languages5: English, French, Spanish, German, Portuguese
Accents28
Voices tagged for narration221
Voices tagged for ads and social228
Voices tagged for customer service256
Voices tagged as characters329
Latency claimSub-50 ms

How does Sume pick a voice?

Sume TTS 1.0 needs a transcript and one voice selector. The documented selectors are the top-level avatar_id or avatar_handle (the voice of one of your workspace avatars, usable when the avatar's voice.status is ready) or voice.id, a Sume TTS voice UUID or a Voices library id that starts with voi_.

A voice name from another vendor is not accepted. The API rejects any other id shape with 400 invalid_voice_id before a job is queued and before credits are reserved, so a wrong voice name costs nothing. If you hold only the API contract, list your avatars at GET /v1/avatar-1.0/avatars and use one whose voice is ready.

  • avatar_id or avatar_handle: resolves the avatar's voice at submit time.
  • voice.id: a TTS voice UUID, or voi_ plus 32 hex characters from your Voices library.
  • If you send an avatar and voice.id together, they must resolve to the same voice, or the request fails with 400.

Which engine says what the voice sounds like?

The engine is separate from the voice. TTS 1.0 has no engine picker and rejects model. To pin an engine, call POST /v1/tts-router/generate with a required model, one of sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview. sonic-latest is documented as an alias for sonic-3.6, never sonic-preview. All of them use the same character price, $0.0475 per 1,000 characters.

Sume TTS selection fields, from the Sume OpenAPI schema (checked 2026-10-10)
FieldWhereRole
avatar_id / avatar_handleTTS 1.0 and TTS RouterPick the voice from a ready avatar
voice.idTTS 1.0 and TTS RouterPick a voice by id
modelTTS Router onlyPick the Sonic engine (required there, rejected on TTS 1.0)
languageBothLanguage the voice speaks the transcript in

When is a bigger catalog the wrong thing to chase?

A catalog of 1,421 voices helps when a person picks a voice by ear once. Automated pipelines do the opposite: they need the same voice on every episode. On Sume the avatar or voice id is that pin, and the job result stores the audio under Sume media URLs you can reuse. Compare that with browsing: the more voices you can pick from, the more you have to document which one your brand uses.

If a named Gradium voice is a hard requirement, Sume is the wrong tool for that step; use Gradium directly, import the finished audio, and use Sume for the parts it does list, such as captions, timeline assembly and lip-sync clips. Imported audio must be public HTTPS or already on the Sume media host depending on the route, so check the route's audio field before you plan the handoff.

Sources

Related posts

More in Models

All Models posts

Written by Sume