Google Chirp 3 Instant Custom Voice: 10 s consent, allowlist only
Google's Chirp 3 Instant Custom Voice is allowlisted, wants a 10-second consent clip, and keeps the cloning key on your side. How that compares to a Sume clone.

Chirp 3 Instant Custom Voice is gated: Google says access is restricted to allow-listed users who contact sales, and the speaker must record a consent statement of up to 10 seconds. The resulting cloning key lives on your side and is sent with each request, per Google's page (read 2026-10-02, page dated 30 September 2026).
On Sume you clone in the app with no allowlist step, but the dialog does not record consent for you, and the API has no clone route.
What Google asks for
The English (US) statement is: "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model." Google wants one single-channel file per recording, with the consent and the reference audio recorded in the same environment.
| Item | What the page says |
|---|---|
| Access | Allow-listed users only |
| Consent clip | Required statement, up to 10 seconds, single channel |
| Encodings | LINEAR16, PCM, MP3, M4A |
| Key storage | Client side, provided per request; no limit on keys |
| Use in synthesis | The voice cloning key goes in the voice_clone parameter |
| Languages | 40+ listed on the page |
How Sume stores a voice
Sume's Voices library in the app has a create dialog that takes a name, a gender, a language and an uploaded or microphone-recorded clip. That dialog has no consent-recording step, and the public API has no route that creates a clone.
A finished Sume voice is a library row with an id, a language and a gender, and text to speech takes the id. The key is not yours to hold, unlike Google's client-side cloning key, so moving a voice to another vendor is not something Sume does for you.
Language note
Sume checks the voice language against the request language. Per the repo's TTS notes, a known mismatch returns HTTP 409 tts_voice_language_mismatch before any job or charge, and you confirm to retry. That protects a cloned voice from being sent text in a language it was not set up for.
Which route to pick
The choice comes down to who must verify the consent.
- Need consent audio verified by the vendor: Google's route, once you have been allow-listed.
- Need a narration voice in the app today and you hold the release yourself: clone in the Voices library.
- Need many languages from one voice: confirm the voice language first, then use the language field on the job.
Sources
Related posts
More in Comparisons
- Flow charges 7 to 15 credits for Omni Flash; Sume bills dollars
Google Flow prices Gemini Omni Flash 720p at 7 to 15 credits by length. Sume bills the same model per second in dollars; clip costs from 3 to 10 seconds.
- Flow video edit costs 40 credits a try; Sume bills per second
Flow charges 40 credits per Gemini Omni video edit. Sume's video_url edit bills output seconds: $0.625 for 5 seconds, $1.25 for 10 at 720p.
- gpt-4o-mini-tts instructions vs Sume TTS speed, volume, emotion
OpenAI steers gpt-4o-mini-tts with free-text instructions. Sume TTS exposes speed 0.6-1.5, volume 0.5-2 and a short emotion string. How the controls compare.
- GPT-6 Sol vs Luna vs Astra: context, prices, cutoffs
GPT-6 Astra, Sol and Luna share a 1,050,000-token window but differ 100x in price. A table from OpenAI's own pages, plus which one Sume runs Formats on.
Written by Sume