Grok custom voices need a live passphrase; Sume takes no uploads

xAI clones a voice from about a minute of live speech after a passphrase check. Sume TTS has no recording-upload field, so here is how to pick a voice instead.

5 min readSume
All posts

How do Grok custom voices work, and can Sume do the same?

Grok custom voices clone your own voice from about a minute of speech recorded live in the xAI console, and Sume does not offer an equivalent: the Sume text-to-speech request has no field for a reference recording. You choose a voice that already exists in your workspace, through an avatar or a voice id.

xAI announced Custom Voices on April 30, 2026. This post reads that page next to the Sume TTS contract so you can decide which side of the line your project sits on. If you need a clone of a specific person, Sume is the wrong tool today and the section on alternatives says what to do instead.

What does xAI actually require to make a custom voice?

The xAI page describes a two-stage check before a voice is created. First the speaker reads a passphrase, and a speech-to-text engine transcribes it and matches it in real time. Second, speaker embeddings from the passphrase and from the full recording are compared to confirm they come from the same person.

xAI says the whole pipeline takes under two minutes and the finished voice works in the Text to Speech endpoints and the Voice Agent API, with the same speech tags, multilingual output and streaming as built-in voices. It also says there is no extra charge for using custom voices in those APIs.

  • About a minute of natural speech, recorded live.
  • Cannot be created from pre-existing recordings.
  • Cannot be used to clone someone else's voice.
  • Usable in Text to Speech and the Voice Agent API.

What can a Sume TTS request select as a voice?

Per the Sume OpenAPI contract, a TTS 1.0 request needs exactly one text input (transcript, 1 to 20,000 characters) plus one voice selector. The selectors are a workspace avatar (avatar_id or avatar_handle, which resolves to that avatar's ready TTS voice) or voice.id.

voice.id must be a TTS voice UUID or a Voices library id beginning voi_. Anything else, such as a voice name from another vendor, is rejected with a 400 invalid_voice_id before a job is queued or credits are reserved. There is no multipart upload, no consent-audio field and no clone mode, and the contract says provider credentials are never sent by the caller.

Voice sourcing, vendor page and Sume OpenAPI (read 2026-10-02)
QuestionGrok custom voicesSume TTS 1.0
Where does the voice come from?Live recording in the xAI consoleAn existing avatar voice, UUID or voi_ id
Reference audio in the API call?Not part of a TTS requestNo field for it
Ownership checkLive passphrase plus speaker-embedding matchNone needed; no cloning step
Cloning someone elseBlocked by the checkNot offered
Clone from an old recordingNot allowedNot offered

When is Sume the right choice anyway?

If the job is narration, a product video or an avatar clip, you usually need a voice that stays the same from line to line, not a copy of a particular person. Pick one avatar voice, reuse its handle on every request, and the voice stays stable across takes. The post on finding a voice id shows the two accepted id shapes.

If the job is "this must sound like me", use a vendor that verifies the speaker, such as xAI above, produce the audio there, and bring the finished file into Sume as a Sume-hosted audio input for timeline work or a talking video. Sume joins and slices audio without re-synthesising it, as described in timeline audio.

Consent matters in both directions. xAI enforces it technically. If you clone a voice anywhere, keep the speaker's written consent with the project files; the voice cloning consent wording post lists examples vendors ask for.

What should you check before choosing?

Decide first whether the audience needs to recognise a specific person. If yes, cloning from a live verified recording is the safer legal footing, and Sume cannot supply it. If no, a stable stock or avatar voice avoids the consent question entirely.

Then confirm language. Sume's language field is a plain BCP-47 or ISO-639 code, and a voice whose primary language differs from the request returns a confirmation step rather than silently speaking. See the mismatch warning post.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume