Voice cloning consent statement example: Azure, Google, OpenAI
Azure, Google and OpenAI each script the consent a speaker records before a clone. Compare the wording, then write your own release for a Sume voice.

A consent statement for voice cloning is a short recording in which the speaker says that they own the voice and agree to a named party making a synthetic version of it. Azure and Google publish exact wording, OpenAI requires one of 16 fixed phrases, and all three check that the consent speaker is the person in the sample. Sume collects none of this, so a Sume clone needs a release you write and keep.
The three wordings
Azure's English (US) text, from Microsoft Learn, names the speaker and the company. Google's, from its Chirp 3 page, is a fixed sentence about owning the voice. OpenAI's guide requires one of its supported phrases, and any divergence fails.
| Vendor | Consent text or rule | Length or format |
|---|---|---|
| Azure personal voice | "I [name] am aware that recordings of my voice will be used by [company] to create and use a synthetic version of my voice." | mp3 or wav, 16 to 48 kHz |
| Google Chirp 3 Instant Custom Voice | "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model." | Up to 10 seconds, single channel |
| OpenAI custom voices | One of 16 supported phrases; any divergence fails | Sample 30 seconds or less |
What they have in common
Those four features are worth copying even when your tool does not enforce them.
- The speaker says it, in their own voice, in the language of the sample.
- The consent is recorded separately from the training sample.
- The vendor compares the consent voice with the sample voice.
- The text names who may use the voice: a company, Google, or the customer.
What a Sume voice needs from you
Sume's Voices library in the app has a create dialog that takes a name, a gender, a language and an uploaded or microphone-recorded clip. That dialog has no consent-recording step, and the public API has no route that creates a clone.
So write your own release. Name the speaker, the project, the length of the licence, where synthetic audio may appear, and how consent can be withdrawn. Record the speaker reading it, and file the audio with the voice id. For the law around AI voices, see the posts on the No Fakes Act and Tennessee's ELVIS Act.
Limits of this post
The wording above is each vendor's own and is for their tools only. It is not legal advice, and a vendor's statement does not satisfy another vendor's check. Re-read the pages before use, since they change. The Sume route is explained in Voice cloning API.
Sources
Related posts
More in Comparisons
- Voxtral TTS 2-minute native limit vs Sume TTS 1200-second cap
Mistral says Voxtral TTS natively generates up to two minutes and the API handles longer. Sume TTS fails audio over 1200 seconds with tts_duration_exceeded.
- Walmart bans seller logos, Coupang wants yours: one prompt per site
Walmart's image guide bars seller logos; Coupang's main-image page says to show yours clearly. Keep a separate prompt per marketplace in a Sume batch.
- WaveSpeed task statuses (timeout, deleted) vs Sume job statuses
WaveSpeed tasks can be created, processing, completed, failed, cancelled, timeout or deleted. Sume has five job statuses. How to map them in a port.
- WaveSpeedAI API alternative: what Sume offers instead
WaveSpeedAI sells 1,000+ models behind one API with tiered rate limits. Sume offers a smaller managed catalog with plan-based concurrency. Compared.
Written by Sume