Gemini voice replication: the 10-30 second clip and consent checklist

Gemini 3.8 voice replication needs a 10-30 second clean clip and a spoken consent statement. A checklist, and why a Gemini voice name will not work on Sume.

4 min readSume
All posts

To replicate a voice with Gemini 3.8 Flash TTS you send two audio files: a clean 10 to 30 second clip of the speaker, and a short consent recording in which the same speaker says Google's required statement word for word. Google's voice replication guide gives the English statement as "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model."

Replication is a feature of the Gemini API, not of Sume. Sume's TTS takes its own voice ids, so a replicated Gemini voice does not carry over. The checklist below is what to prepare on the Google side and what to keep on yours.

What the Google docs specify

The speech generation guide says voices are created with POST /v1beta/voices, with a type of prompted for voice design from a text description and replicated for a voice built from reference audio. The replication guide covers the audio. Both pages were read on 2026-10-07.

Gemini voice replication requirements, from Google's guides (read 2026-10-07)
ItemWhat Google states
Reference clip10 to 30 seconds of clean, natural speech from the speaker to replicate
Consent recordingThe speaker clearly recites the mandatory consent statement in one of the supported languages
Consent statement (English)"I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model."
Recommended format24 kHz mono 16-bit WAV; both recordings from the same adult speaker
Consent languages30 supported language locales for the phrase
Voice creation endpointPOST /v1beta/voices with type replicated

Prepare the clip and the consent file separately

Treat the two recordings as two deliverables. The clip is for the model, so it should be one speaker, no music bed, no cross-talk, and no room echo you do not want copied into every future line. The consent file is for your records, so it should be a clean take of the statement, recorded by the person whose voice it is, and stored with the date.

  • Record the clip in the register you will use: an ad voice, a calm explainer voice, not a one-off laugh.
  • Use the statement text exactly as the guide gives it for your language; do not paraphrase it.
  • Ask the speaker to record the consent file themselves, in the same session if you can, so the match is easy to defend.
  • Write down what the voice may be used for before you upload anything: which brand, which channels, until when.

What does not transfer to Sume

Sume's TTS 1.0 selects a voice by an avatar reference, or by voice.id. The OpenAPI contract says that id must be a TTS voice UUID or a Voices library id (voi_ plus 32 hex characters), and that a value of any other shape, such as a voice name from another TTS ecosystem, is rejected with HTTP 400 invalid_voice_id before a job is queued or credits are reserved.

In practice that means a Gemini voice name or a Gemini voice reference is not a Sume voice. If you want the same person's voice on a Sume job, that is a separate decision with its own consent, made through Sume's own voice routes, and not something this post can promise on Google's behalf.

Records to keep either way

Whichever engine speaks the final lines, a short file next to the project saves an argument later.

  • The consent recording and its date, kept with the reference clip.
  • The Google voice resource name and the project it was created in, plus the date you plan to re-create it.
  • The scripts that were voiced and the output files, so you can show what the voice said.
  • The platform where each file was published and whether you disclosed AI voice there.

Sources

Related posts

More in Models

All Models posts

Written by Sume