Gemini TTS has 30 prebuilt voices; Sume TTS takes avatar voice ids
Google's Gemini TTS page lists 30 prebuilt voices. Sume TTS takes a voice UUID or voi_ id, or an avatar whose voice.status is ready; cloning stays app-only.
Google's Gemini speech page lists 30 prebuilt voices, while Sume TTS does not offer a menu of named stock voices. You send either a voice UUID or a voi_ id, or you point at an avatar whose voice.status is ready and let Sume resolve its voice; voice cloning is available in the Sume app only, not through the API.
Google facts are from its page Text-to-speech generation (TTS) | Gemini API (read 2026-10-10). Sume facts are from the Sume OpenAPI contract and the Avatar docs.
How the two pick a voice
Gemini's page lists prebuilt voices such as Kore, Puck and Charon, and says its Flash TTS model supports over 130 languages. It also describes stored custom voices: 200 per project with a one-year lifetime for stateful voices, and a seven-day lifetime for stateless voice keys. Multi-speaker output supports up to two speakers using prebuilt voices. Sume's contract describes one selector that a caller can discover without a separate catalog.
| Question | Gemini TTS | Sume TTS 1.0 |
|---|---|---|
| Named stock voices | 30 prebuilt | None; other names return 400 invalid_voice_id |
| Selector | Voice name | voice.id (UUID or voi_ id), avatar_id or avatar_handle |
| Custom voices | Stored, 200 per project | Created in the Sume app; API cannot clone |
| Default output | WAV, or raw PCM when streaming | MP3 44.1 kHz at 128 kbps |
The avatar route in practice
GET /v1/avatar-1.0/avatars lists workspace avatars. Any avatar whose voice.status is ready can be used by passing its avatar_id or avatar_handle on the TTS request, and Sume resolves the voice when the job is submitted. If you also send voice.id, the two must match or the request fails with 400. This is the way a caller who holds only the contract finds a usable voice. For avatar lip-sync muxing, the contract suggests asking for WAV with pcm_s16le at 44,100 Hz.
Language and length differences
Sume's language field takes a BCP-47 or ISO-639 code, and a non-English transcript should always set it. I do not give a Sume language count, because I did not verify one from a Sume source. Input length is capped at 20,000 characters and generated audio at 1,200 seconds. I did not read a comparable limit in the Google page, so I draw no comparison.
Which should you pick
If you want to browse a fixed set of voices by name, Gemini fits better. If your voices live in a Sume workspace, such as a brand avatar, Sume's selector keeps one voice across video, lip-sync and voice-over without a second vendor.
Sources
Related posts
More in Comparisons
- Grok Imagine's 4 keyframes and 7 references vs Sume's one image
xAI's Grok Imagine 1.5 takes up to 4 keyframes and up to 7 references. Sume's grok-imagine-video-1.5 row takes one image only. Rows to use for multi-image work.
- Grok Imagine Lite upscales to 1080p: what Sume offers instead
xAI describes Grok Imagine Video 1.5 Lite as lowest cost with upscaled 1080p. Sume has no Lite row; it lists grok-imagine-video-1.5 at a flat rate. Compare.
- Grok Imagine's 3 voice references vs Sume reference-audio rows
xAI's Grok Imagine 1.5 takes up to 3 voice references. Sume's Grok row takes none; Seedance, Wan 3.0 and MiniMax accept reference audio under limits.
- Grok Imagine Lite draft plus Sume upscale: what a 10 s clip totals
xAI lists Lite at $0.020 per second and 1.5 at $0.080. Add Sume video upscale at $0.009 per input second and a 10 s clip totals $0.29. Read 2026-10-10.
Written by Sume