Gradium's 329 character voices vs voices of Sume avatars
Gradium lists 329 character voices. On Sume, a spoken character is an avatar with a ready voice; Avatar 1.0 is English-only. What that means for casting.
On Sume, a character that talks is an avatar: you create it once, it gets a voice, and the avatar handle then drives both the face and the speech. Avatar 1.0 is English-only, so Gradium's French, Spanish, German and Portuguese character voices have no equivalent there. Gradium's page lists 329 voices tagged as characters among its 1,421 (read 2026-10-10 at Gradium).
This post is for anyone choosing between a big voice catalog and the avatar route for a recurring character, using Sume's avatar video guide, avatar creation guide and the API reference.
How does an avatar get a voice on Sume?
The Sume schema says TTS can select a voice with avatar_id or avatar_handle: Sume loads the avatar in your workspace and uses its TTS voice id. That works exactly when the avatar's voice.status is ready, which you can read from GET /v1/avatar-1.0/avatars. The same schema notes that selecting a voice this way does not change the talking-head pipeline; it only picks a voice for TTS.
Voice cloning from your own recording is a feature of the Sume app, not of the API. Over the API you reference existing avatars and voice ids.
What does a talking avatar video need?
The avatar-video route takes avatar_handle plus a script or multi-scene video_inputs. Scripts and plans must estimate to 4 to 60 seconds. Aspect ratios are 1:1, 3:4, 9:16, 4:3 and 16:9 (default 9:16), and resolution is 720p. Quality is plus by default, with standard the fastest path and max the highest tier. A voice.type: "silence" beat gives a pause with no speech and needs an explicit duration.
| Item | Value |
|---|---|
| Script length window | 4 to 60 seconds, estimated |
| Aspect ratios | 1:1, 3:4, 9:16, 4:3, 16:9 |
| Resolution | 720p |
| Quality tiers | standard, plus (default), max |
| Language | English only (Avatar 1.0) |
When should a character be a TTS voice instead?
If you only need a narrator over footage, and not a face, TTS with a voice id is cheaper and simpler: no avatar, no video render. If the character has a face on screen, use the avatar route so the mouth follows the speech. Both can reuse the same voice, because the avatar's voice is what TTS resolves.
- Face on screen, English: avatar video with
avatar_handle. - Voice only: TTS with the avatar's voice or a voice id.
- Non-English character voice: not available through Avatar 1.0; produce that audio elsewhere.
What about a casting catalog like Gradium's?
A large tagged catalog saves time when you audition many voices. Sume's route is different: decide the character once, create the avatar, and pin it. That fits series work, such as a recurring host, better than an audition-first flow. If your characters need five languages, the honest answer is that Sume's avatar path does not cover that today.
What does casting look like in practice?
Start with the language. If the character must speak anything other than English, Avatar 1.0 is out, and the voice-only route is the one to use. If English is enough, write a short brief for the character: age feel, pace, warmth, the kind of lines it will read. Then create the avatar in the Sume app, wait for the voice to reach ready, and run a 4 to 10 second test script through the avatar video route before building a series on it.
Keep the script natural for speech. The duration check is an estimate, so a script near the 4 or 60 second edges may fall outside the window; aim for the middle of the range for reliable results. For a longer piece, split it into multi-scene video_inputs rather than one long script.
- Test with one short script before committing to a series.
- Pin the avatar handle in your config, not a display name.
- Re-check
voice.statusbefore every batch.
Sources
Related posts
More in Sume Avatar 1.0
- Moving Company Spokesperson Ad: A 20-Second Avatar for About $4.90
A 20-second plus-quality Avatar 1.0 spokesperson clip costs about $4.90 on Sume, at 98 cents per 4 seconds. English only, with captions at $0.20 extra.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
Written by Sume