Gemini 3.8 Flash TTS voice library vs Sume voice selectors

Gemini 3.8 Flash TTS lists 30 curated voices plus a larger extended library. Sume picks a voice by avatar or voice id. How the two selection models differ.

4 min readSume
All posts

Gemini 3.8 Flash TTS gives you a browsable catalog: 30 curated prebuilt voices, hundreds more in an Extended Voice Library, and custom voices you design or replicate (Gemini docs, read 2026-10-07). The Sume TTS 1.0 request has no browse-and-audition step. You name a voice through an avatar or a voice id, and Sume resolves it at submit time.

How Gemini 3.8 Flash TTS picks a voice

The Gemini docs describe three tiers. The prebuilt studio collection has 30 curated voices. An Extended Voice Library adds hundreds more. Custom voices come from voice design (a text prompt that yields a voice_... id) or voice replication (a replicated or stateless voicekey_...).

Storage has limits worth planning around: a project can hold at most 200 stored custom voices, with one-year retention, and stateless voice keys expire after 7 days.

Voice selection limits on Gemini TTS (read 2026-10-07)
ItemGemini 3.8 Flash TTS
Curated prebuilt voices30
Extended libraryHundreds more
Stored custom voices per project200 (1-year retention)
Stateless voice keysExpire in 7 days
Speakers in one requestUp to 2 prebuilt voices

How Sume TTS 1.0 picks a voice

A Sume TTS request carries one voice selector. The discoverable one is the avatar reference: list GET /v1/avatar-1.0/avatars, take any avatar whose voice.status is ready, and send avatar_id or avatar_handle. If you already hold a Sume TTS voice id, send voice.id instead. If you send both, they must match or the request fails with 400.

The request also takes a language. A known mismatch between the voice's primary language and the requested language returns HTTP 409 tts_voice_language_mismatch before any job or charge, and you retry with confirm_language_mismatch: true only after a person agrees.

Which model fits which job

If your decision is "audition 200 voices and choose one by ear," Gemini's library is the stronger fit today, and the TTS 1.0 request does not offer a comparable browse step. If your decision is "the same speaker voices every video for this brand," the Sume avatar-to-voice binding keeps the choice in one place, so each job only needs the avatar handle.

Sume TTS Router is a separate surface that adds a required model field. It lists Cartesia Sonic ids only (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). It does not route to Gemini TTS, so a Gemini voice cannot be called through Sume.

  • Pick a Gemini voice if you want the largest voice menu and are fine with a separate Google account.
  • Pick a Sume avatar voice if the voice should stay tied to a presenter across TTS and avatar video.
  • Do not mix the two ids: a Gemini voice_... id means nothing to Sume.

A request that names the voice

This is the shape from the Sume OpenAPI examples, with an avatar selector and a language.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: voice-pick-001" \
  -d '{
    "transcript": "Hello from this avatar.",
    "avatar_handle": "speaker",
    "language": "en",
    "mode": "async"
  }'

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume