Grok's 26 voices across 25+ languages vs how Sume picks a voice
xAI lists 26 Grok voices for support, characters, commentary, ads and education. Sume TTS has no voice-name catalog; it takes an avatar or a voice id.

How many voices does Grok have, and does Sume match them?
xAI says 21 new flagship voices joined the original five on July 6, 2026, with the five originals improved, for 26 in total. Sume does not publish a voice-name list in its TTS contract, so you cannot ask for a voice called "Atlas"; you select a voice by avatar or by id.
This matters when you port a script. A request that names a vendor voice will not work on Sume, and Sume returns a clear error before charging anything.
What does the xAI voice release say?
The announcement says each voice was cast for a specific job: support, characters, commentary, advertising and education. The voices are available in the realtime Voice Agent API, the Text to Speech API and the new Grok Voice Agent Builder in the xAI console, and the page lists 25+ languages.
Named voices on the page include Naksh, Atlas, Aurora, Liora, Carina, Zagan, Helix, Orion, Luna and Wellness. The TTS docs add that you can list voices programmatically, and that eve is the default.
How does voice selection work on Sume?
The Sume TTS 1.0 contract offers two routes. The first is an avatar: send avatar_id or avatar_handle, from GET /v1/avatar-1.0/avatars, and Sume resolves that avatar's TTS voice, which works when the avatar's voice status is ready. The second is voice.id, which must be a TTS voice UUID or a Voices library id starting voi_.
If you send both, they must resolve to the same voice or the request fails with 400. Names from other ecosystems are rejected synchronously with invalid_voice_id.
| Aspect | xAI | Sume TTS 1.0 |
|---|---|---|
| Pick by | Voice name such as eve | avatar_id, avatar_handle or voice.id |
| Discover voices | List voices programmatically | GET /v1/avatar-1.0/avatars, check voice.status |
| Language per voice | All voices speak every supported language | language field you set; mismatch needs confirmation |
| Wrong id shape | Not covered on the pages read | 400 invalid_voice_id, nothing reserved |
Which job should each voice do?
xAI casts voices by job, which is a sensible way to choose even when the catalog is not portable. Use the same five jobs as a checklist and decide what each needs before you audition anything.
- Support: calm, clear, steady pace; test with numbers, order ids and addresses.
- Characters: strong personality; expect to keep one voice per character for the whole project.
- Commentary: faster, energetic delivery; set speed in generation_config rather than rewriting the script.
- Advertising: short reads with brand names; use pronunciation_dict_id for names that mis-read.
- Education: neutral and warm; check long passages for drift in tone.
What should you do when moving a Grok script to Sume?
Map each Grok voice to a Sume avatar voice by listening, not by name. Generate the same two sentences with each candidate and keep the closest. Record the avatar handle in your project config, so every later render reuses it.
Then set language explicitly. Sume treats a missing language as English at the provider, with a Hangul or kana-only fallback for Korean and Japanese, so a Spanish script without language: "es" will be read as English. The language default post shows the failure.
Finally check pricing before a batch. The page for xAI does not list a price, so compare against the Grok per-million-character quote and use the Sume dry-run cost preview on your own script length.
Sources
Related posts
More in Comparisons
- Grok TTS takes 60,000 characters; Sume takes 20,000: how to split
Grok TTS takes 60,000 characters per request; Sume TTS takes 20,000. How to split a long script into jobs and join the audio without gaps.
- Grok TTS codecs and sample rates vs Sume TTS output_format
xAI TTS and Sume TTS offer the same six sample rates and mp3 bit-rate range. They differ on defaults (24 kHz vs 44.1 kHz) and Sume adds raw and float PCM.
- Grok TTS language "auto" vs Sume's explicit language field
xAI TTS can auto-detect the language of your text. Sume TTS cannot: omit language and Spanish is read as English, except Korean and Japanese.
- Grok TTS speech tags like laugh and whisper vs Sume's emotion field
xAI lets you write [laugh] and wrap text in whisper or singing tags. Sume TTS documents no tag grammar, only emotion, speed and volume in generation_config.
Written by Sume