Voice replication API audit checklist before you switch
Gemini 3.8 Flash TTS is GA with voice replication and 150+ voices. Before switching providers, audit these items against Sume's live catalog.

The Gemini API changelog for Sep 22, 2026 says Gemini 3.8 Flash TTS and Flash-Lite TTS are generally available, with voice design, voice replication and "150+ prebuilt and custom voices". Before you move a voice workload to Sume, check the live catalog for each capability below; the docs I read do not describe voice replication.
What Google lists
| Item | Listed on Sep 22, 2026 |
|---|---|
| Models | Gemini 3.8 Flash TTS and Flash-Lite TTS |
| Status | Generally available |
| Features | Voice design and voice replication |
| Voices | 150+ prebuilt and custom voices |
What Sume docs confirm
The hosted MCP lists a paid tts_create tool, with tts_source_get and tts_source_verify_spine to check accepted scripts against generated jobs. The models overview says on-camera speaking shots are made from an accepted still plus TTS, because video models do not lip-sync to generated TTS. It does not list a voice-replication endpoint.
The audit
Run these checks against the live catalog and your current output. Mark each one pass, fail or unknown before you commit.
- Does the catalog list the voice or voice id you use today?
- Can you create or upload a custom voice, and who must consent?
- Which languages and accents does each voice cover?
- Is output a hosted URL, and how long does it stay available?
- What is the price unit, per character, per second or per job?
- How does the API behave on long scripts: one call or split?
- Can you re-run the same script and get the same voice?
- What does a failure look like, and is it billed?
Decide on evidence
If voice replication is the reason you are switching, and Sume's catalog does not list it, stay where you are. Listen to a short sample from each provider before you decide.
Sources
Related posts
More in Developers
- Reuse TTS word timestamps as caption words, skip a second STT
You already know what the voice said and when. Feed the TTS word timings to the caption job as `words` so brand names are never misheard by speech-to-text.
- A "use step" function that submits a Sume job once
Put the Sume submit call inside one "use step" function and derive the Idempotency-Key from a stable run key, so a retried step returns the original job.
- Veo is US-region only: what EU teams call on Sume
A Sept 2026 listing says Veo runs only in us-central1, with no EU region documented. Sume has one video endpoint and its docs state no region.
- Vercel renamed Edge Requests: match Sume errors by code, not text
Vercel Edge Requests are now CDN Requests, and billing reads should use SkuId. Same rule for Sume errors: branch on error.code and request_id, never message.
Written by Sume