Gemini custom voices: 200 per project, 1-year expiry, 7-day voice keys

Gemini 3.8 TTS stores up to 200 prompted or replicated voices per project. Voices expire after a year and voice keys after 7 days. What to plan for.

4 min readSume
All posts

A Google project can hold up to 200 custom voices for Gemini 3.8 TTS, counting prompted and replicated voices together. Per Google's speech generation guide, a stored voice has a one-year time to live, and the voicekey_ keys it describes have a 7-day TTL.

So a custom voice is a rental, not an asset. If a brand voice has to sound the same in 18 months, plan to re-create it, and keep the inputs you need to do that.

The limits in one table

Voices are created with POST /v1beta/voices. The type decides what you supply.

Gemini custom voice limits, from Google's speech and voice replication guides (read 2026-10-07)
ItemValue
Voice typesprompted (from a text description) and replicated (from reference audio)
Voices per project200, prompted and replicated combined
Voice time to live1 year
voicekey_ key time to live7 days
Replication inputs10 to 30 second reference clip plus a consent recording
Modelsgemini-3.8-flash-tts and gemini-3.8-flash-lite-tts

Plan the lifecycle

Treat the voice resource like a certificate: it has an issue date and an expiry, and somebody owns the renewal.

  • Record the creation date and add a reminder well before the one-year mark.
  • For a prompted voice, store the exact description text you used, since that is what you re-submit.
  • For a replicated voice, keep the reference clip and the consent recording together, because a re-creation needs both.
  • Do not cache a voice key for longer than its 7-day life; fetch a fresh one in the job that needs it.
  • Count voices per project. Two hundred is plenty for one brand and tight for an agency that makes one voice per client.

A renewal checklist

Put these on a calendar for each custom voice you create, so an expiry is a planned task rather than a surprise on launch day.

  • Creation date, expiry date and the owner of the renewal.
  • The type of voice, prompted or replicated, and the inputs needed to rebuild it.
  • The campaigns that depend on the voice, so you know what to re-render if the sound shifts slightly.
  • A rendered reference line saved as a file, so you can compare the rebuilt voice against the original by ear.

Why this matters for a mixed pipeline

A voice that lives inside Google's project is only addressable from Google's API. Sume's TTS 1.0 takes a TTS voice UUID or a voi_ id, or an avatar reference, and answers any other voice string with HTTP 400 invalid_voice_id before a job is queued, so there is no way to point a Sume job at a Gemini voice resource.

If a project needs one consistent voice across tools, pick the engine that will outlive the project, and render the final lines to audio files you own. Files do not expire after a year. Sume's timeline audio route can then join or trim those Sume-hosted files at $0.01 per job.

Sources

Related posts

More in Models

All Models posts

Written by Sume