Gemini TTS has 30 prebuilt voices; Sume TTS takes avatar voice ids

Google's Gemini TTS page lists 30 prebuilt voices. Sume TTS takes a voice UUID or voi_ id, or an avatar whose voice.status is ready; cloning stays app-only.

5 min readSume
All posts

Google's Gemini speech page lists 30 prebuilt voices, while Sume TTS does not offer a menu of named stock voices. You send either a voice UUID or a voi_ id, or you point at an avatar whose voice.status is ready and let Sume resolve its voice; voice cloning is available in the Sume app only, not through the API.

Google facts are from its page Text-to-speech generation (TTS) | Gemini API (read 2026-10-10). Sume facts are from the Sume OpenAPI contract and the Avatar docs.

How the two pick a voice

Gemini's page lists prebuilt voices such as Kore, Puck and Charon, and says its Flash TTS model supports over 130 languages. It also describes stored custom voices: 200 per project with a one-year lifetime for stateful voices, and a seven-day lifetime for stateless voice keys. Multi-speaker output supports up to two speakers using prebuilt voices. Sume's contract describes one selector that a caller can discover without a separate catalog.

Voice selection (Google read 2026-10-10; Sume per OpenAPI)
QuestionGemini TTSSume TTS 1.0
Named stock voices30 prebuiltNone; other names return 400 invalid_voice_id
SelectorVoice namevoice.id (UUID or voi_ id), avatar_id or avatar_handle
Custom voicesStored, 200 per projectCreated in the Sume app; API cannot clone
Default outputWAV, or raw PCM when streamingMP3 44.1 kHz at 128 kbps

The avatar route in practice

GET /v1/avatar-1.0/avatars lists workspace avatars. Any avatar whose voice.status is ready can be used by passing its avatar_id or avatar_handle on the TTS request, and Sume resolves the voice when the job is submitted. If you also send voice.id, the two must match or the request fails with 400. This is the way a caller who holds only the contract finds a usable voice. For avatar lip-sync muxing, the contract suggests asking for WAV with pcm_s16le at 44,100 Hz.

Language and length differences

Sume's language field takes a BCP-47 or ISO-639 code, and a non-English transcript should always set it. I do not give a Sume language count, because I did not verify one from a Sume source. Input length is capped at 20,000 characters and generated audio at 1,200 seconds. I did not read a comparable limit in the Google page, so I draw no comparison.

Which should you pick

If you want to browse a fixed set of voices by name, Gemini fits better. If your voices live in a Sume workspace, such as a brand avatar, Sume's selector keeps one voice across video, lip-sync and voice-over without a second vendor.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume