Grok custom voices need a live passphrase; Sume takes no uploads
xAI clones a voice from about a minute of live speech after a passphrase check. Sume TTS has no recording-upload field, so here is how to pick a voice instead.

How do Grok custom voices work, and can Sume do the same?
Grok custom voices clone your own voice from about a minute of speech recorded live in the xAI console, and Sume does not offer an equivalent: the Sume text-to-speech request has no field for a reference recording. You choose a voice that already exists in your workspace, through an avatar or a voice id.
xAI announced Custom Voices on April 30, 2026. This post reads that page next to the Sume TTS contract so you can decide which side of the line your project sits on. If you need a clone of a specific person, Sume is the wrong tool today and the section on alternatives says what to do instead.
What does xAI actually require to make a custom voice?
The xAI page describes a two-stage check before a voice is created. First the speaker reads a passphrase, and a speech-to-text engine transcribes it and matches it in real time. Second, speaker embeddings from the passphrase and from the full recording are compared to confirm they come from the same person.
xAI says the whole pipeline takes under two minutes and the finished voice works in the Text to Speech endpoints and the Voice Agent API, with the same speech tags, multilingual output and streaming as built-in voices. It also says there is no extra charge for using custom voices in those APIs.
- About a minute of natural speech, recorded live.
- Cannot be created from pre-existing recordings.
- Cannot be used to clone someone else's voice.
- Usable in Text to Speech and the Voice Agent API.
What can a Sume TTS request select as a voice?
Per the Sume OpenAPI contract, a TTS 1.0 request needs exactly one text input (transcript, 1 to 20,000 characters) plus one voice selector. The selectors are a workspace avatar (avatar_id or avatar_handle, which resolves to that avatar's ready TTS voice) or voice.id.
voice.id must be a TTS voice UUID or a Voices library id beginning voi_. Anything else, such as a voice name from another vendor, is rejected with a 400 invalid_voice_id before a job is queued or credits are reserved. There is no multipart upload, no consent-audio field and no clone mode, and the contract says provider credentials are never sent by the caller.
| Question | Grok custom voices | Sume TTS 1.0 |
|---|---|---|
| Where does the voice come from? | Live recording in the xAI console | An existing avatar voice, UUID or voi_ id |
| Reference audio in the API call? | Not part of a TTS request | No field for it |
| Ownership check | Live passphrase plus speaker-embedding match | None needed; no cloning step |
| Cloning someone else | Blocked by the check | Not offered |
| Clone from an old recording | Not allowed | Not offered |
When is Sume the right choice anyway?
If the job is narration, a product video or an avatar clip, you usually need a voice that stays the same from line to line, not a copy of a particular person. Pick one avatar voice, reuse its handle on every request, and the voice stays stable across takes. The post on finding a voice id shows the two accepted id shapes.
If the job is "this must sound like me", use a vendor that verifies the speaker, such as xAI above, produce the audio there, and bring the finished file into Sume as a Sume-hosted audio input for timeline work or a talking video. Sume joins and slices audio without re-synthesising it, as described in timeline audio.
Consent matters in both directions. xAI enforces it technically. If you clone a voice anywhere, keep the speaker's written consent with the project files; the voice cloning consent wording post lists examples vendors ask for.
What should you check before choosing?
Decide first whether the audience needs to recognise a specific person. If yes, cloning from a live verified recording is the safer legal footing, and Sume cannot supply it. If no, a stable stock or avatar voice avoids the consent question entirely.
Then confirm language. Sume's language field is a plain BCP-47 or ISO-639 code, and a voice whose primary language differs from the request returns a confirmation step rather than silently speaking. See the mismatch warning post.
Sources
Related posts
More in Comparisons
- Grok's 26 voices across 25+ languages vs how Sume picks a voice
xAI lists 26 Grok voices for support, characters, commentary, ads and education. Sume TTS has no voice-name catalog; it takes an avatar or a voice id.
- Grok TTS takes 60,000 characters; Sume takes 20,000: how to split
Grok TTS takes 60,000 characters per request; Sume TTS takes 20,000. How to split a long script into jobs and join the audio without gaps.
- Grok TTS codecs and sample rates vs Sume TTS output_format
xAI TTS and Sume TTS offer the same six sample rates and mp3 bit-rate range. They differ on defaults (24 kHz vs 44.1 kHz) and Sume adds raw and float PCM.
- Grok TTS language "auto" vs Sume's explicit language field
xAI TTS can auto-detect the language of your text. Sume TTS cannot: omit language and Spanish is read as English, except Korean and Japanese.
Written by Sume