Gradium's 329 character voices vs voices of Sume avatars

Gradium lists 329 character voices. On Sume, a spoken character is an avatar with a ready voice; Avatar 1.0 is English-only. What that means for casting.

5 min readSume
All posts

On Sume, a character that talks is an avatar: you create it once, it gets a voice, and the avatar handle then drives both the face and the speech. Avatar 1.0 is English-only, so Gradium's French, Spanish, German and Portuguese character voices have no equivalent there. Gradium's page lists 329 voices tagged as characters among its 1,421 (read 2026-10-10 at Gradium).

This post is for anyone choosing between a big voice catalog and the avatar route for a recurring character, using Sume's avatar video guide, avatar creation guide and the API reference.

How does an avatar get a voice on Sume?

The Sume schema says TTS can select a voice with avatar_id or avatar_handle: Sume loads the avatar in your workspace and uses its TTS voice id. That works exactly when the avatar's voice.status is ready, which you can read from GET /v1/avatar-1.0/avatars. The same schema notes that selecting a voice this way does not change the talking-head pipeline; it only picks a voice for TTS.

Voice cloning from your own recording is a feature of the Sume app, not of the API. Over the API you reference existing avatars and voice ids.

What does a talking avatar video need?

The avatar-video route takes avatar_handle plus a script or multi-scene video_inputs. Scripts and plans must estimate to 4 to 60 seconds. Aspect ratios are 1:1, 3:4, 9:16, 4:3 and 16:9 (default 9:16), and resolution is 720p. Quality is plus by default, with standard the fastest path and max the highest tier. A voice.type: "silence" beat gives a pause with no speech and needs an explicit duration.

Avatar video facts from Sume's docs (checked 2026-10-10)
ItemValue
Script length window4 to 60 seconds, estimated
Aspect ratios1:1, 3:4, 9:16, 4:3, 16:9
Resolution720p
Quality tiersstandard, plus (default), max
LanguageEnglish only (Avatar 1.0)

When should a character be a TTS voice instead?

If you only need a narrator over footage, and not a face, TTS with a voice id is cheaper and simpler: no avatar, no video render. If the character has a face on screen, use the avatar route so the mouth follows the speech. Both can reuse the same voice, because the avatar's voice is what TTS resolves.

  • Face on screen, English: avatar video with avatar_handle.
  • Voice only: TTS with the avatar's voice or a voice id.
  • Non-English character voice: not available through Avatar 1.0; produce that audio elsewhere.

What about a casting catalog like Gradium's?

A large tagged catalog saves time when you audition many voices. Sume's route is different: decide the character once, create the avatar, and pin it. That fits series work, such as a recurring host, better than an audition-first flow. If your characters need five languages, the honest answer is that Sume's avatar path does not cover that today.

What does casting look like in practice?

Start with the language. If the character must speak anything other than English, Avatar 1.0 is out, and the voice-only route is the one to use. If English is enough, write a short brief for the character: age feel, pace, warmth, the kind of lines it will read. Then create the avatar in the Sume app, wait for the voice to reach ready, and run a 4 to 10 second test script through the avatar video route before building a series on it.

Keep the script natural for speech. The duration check is an estimate, so a script near the 4 or 60 second edges may fall outside the window; aim for the middle of the range for reliable results. For a longer piece, split it into multi-scene video_inputs rather than one long script.

  • Test with one short script before committing to a series.
  • Pin the avatar handle in your config, not a display name.
  • Re-check voice.status before every batch.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume