Gemini 3.8 Flash TTS voice library vs Sume voice selectors
Gemini 3.8 Flash TTS lists 30 curated voices plus a larger extended library. Sume picks a voice by avatar or voice id. How the two selection models differ.

Gemini 3.8 Flash TTS gives you a browsable catalog: 30 curated prebuilt voices, hundreds more in an Extended Voice Library, and custom voices you design or replicate (Gemini docs, read 2026-10-07). The Sume TTS 1.0 request has no browse-and-audition step. You name a voice through an avatar or a voice id, and Sume resolves it at submit time.
How Gemini 3.8 Flash TTS picks a voice
The Gemini docs describe three tiers. The prebuilt studio collection has 30 curated voices. An Extended Voice Library adds hundreds more. Custom voices come from voice design (a text prompt that yields a voice_... id) or voice replication (a replicated or stateless voicekey_...).
Storage has limits worth planning around: a project can hold at most 200 stored custom voices, with one-year retention, and stateless voice keys expire after 7 days.
| Item | Gemini 3.8 Flash TTS |
|---|---|
| Curated prebuilt voices | 30 |
| Extended library | Hundreds more |
| Stored custom voices per project | 200 (1-year retention) |
| Stateless voice keys | Expire in 7 days |
| Speakers in one request | Up to 2 prebuilt voices |
How Sume TTS 1.0 picks a voice
A Sume TTS request carries one voice selector. The discoverable one is the avatar reference: list GET /v1/avatar-1.0/avatars, take any avatar whose voice.status is ready, and send avatar_id or avatar_handle. If you already hold a Sume TTS voice id, send voice.id instead. If you send both, they must match or the request fails with 400.
The request also takes a language. A known mismatch between the voice's primary language and the requested language returns HTTP 409 tts_voice_language_mismatch before any job or charge, and you retry with confirm_language_mismatch: true only after a person agrees.
Which model fits which job
If your decision is "audition 200 voices and choose one by ear," Gemini's library is the stronger fit today, and the TTS 1.0 request does not offer a comparable browse step. If your decision is "the same speaker voices every video for this brand," the Sume avatar-to-voice binding keeps the choice in one place, so each job only needs the avatar handle.
Sume TTS Router is a separate surface that adds a required model field. It lists Cartesia Sonic ids only (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). It does not route to Gemini TTS, so a Gemini voice cannot be called through Sume.
- Pick a Gemini voice if you want the largest voice menu and are fine with a separate Google account.
- Pick a Sume avatar voice if the voice should stay tied to a presenter across TTS and avatar video.
- Do not mix the two ids: a Gemini
voice_...id means nothing to Sume.
A request that names the voice
This is the shape from the Sume OpenAPI examples, with an avatar selector and a language.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: voice-pick-001" \
-d '{
"transcript": "Hello from this avatar.",
"avatar_handle": "speaker",
"language": "en",
"mode": "async"
}'Sources
Related posts
More in Comparisons
- Gemini API leads with Omni Flash, Veo 3.1 for specialists: and Sume?
Google's video docs recommend Gemini Omni Flash as the default and Veo 3.1 for specific needs. Sume has sume/auto or a pinned catalog id. How they line up.
- Grok Imagine, Qwen Image and Imagen 4 Fast all cost 2.5 cents on Sume
Grok Imagine, Qwen Image and Imagen 4 Fast each bill $0.025 per image on Sume. What separates them: ratios, n, edits. Choose the cheap row that fits your job.
- HappyHorse 1.0 vs 1.1: which one renders 480P on Model Studio?
On Alibaba Model Studio only HappyHorse 1.1 lists 480P. Both versions take 3 to 15 seconds, and 1.0 adds a video-edit id. Neither is in Sume's catalog.
- Headshot background to neutral grey: cutout plus Pillow or an AI edit
Swap a headshot background for grey: Sume RMBG at $0.0225 plus a Pillow composite keeps the face pixels exact; an Ideogram 4.5 edit costs $0.075 or more.
Written by Sume