Azure personal voice consent: name and company must match audio
Azure personal voice requires a recorded consent statement whose talent name and company match the audio. The fields, formats and what Sume collects instead.

Azure personal voice makes consent a structured record: a recorded statement in which the voice talent says their own name and the company's name, uploaded before any training audio. Per Microsoft's consent page (read 2026-10-02; page metadata shows an update on 2026-09-16), the name and company you enter must match what was spoken, and neither can be changed later.
Sume stores no such record. Its clone dialog takes a clip, a name, a gender and a language, so the consent proof stays with you.
What the consent step contains
The English (US) statement is: "I [state your first and last name] am aware that recordings of my voice will be used by [state the name of the company] to create and use a synthetic version of my voice." The language of the statement must match the language of the training data, and Microsoft says the statement is used to check that the talent is the same person as the speaker in the audio prompt.
| Field or rule | Detail |
|---|---|
| Required properties | projectId, voiceTalentName, companyName, locale, and the audio file |
| Immutable after creation | voiceTalentName, companyName, locale, and the consent id |
| Upload path | POST with audiodata, or PUT with an audioUrl (SAS URL) |
| Audio formats | mp3 or wav at 16, 24, 44.1 or 48 kHz |
| Locale rule | Consent language must match the training data language |
Where Sume differs
Sume has no consent object, no talent name field and no company field. Sume's Voices library in the app has a create dialog that takes a name, a gender, a language and an uploaded or microphone-recorded clip. That dialog has no consent-recording step, and the public API has no route that creates a clone.
Because the API cannot create a clone, there is also nothing to call with a consent id. You pass a finished voice id to text to speech.
What to copy from Azure's design
Even if you never use Azure, its consent design is a good template for your own release.
- Record the consent in the same language as the sample.
- Name the person and the company in the spoken statement, then keep both strings in your release file.
- Treat the name and company as fixed: if either changes, record a new statement.
- Keep the consent audio with the voice id in your own storage, since Sume does not attach it.
Next step
If your vendor requires consent audio, use that vendor's route for that voice. If you only need narration in a voice you are licensed to use, clone it in the app and use the id in text to speech. A longer view of what Sume does and does not offer is in Voice cloning API.
Sources
Related posts
More in Comparisons
- Bannerbear 60 POSTs per 10 seconds vs Sume's per-minute budgets
Bannerbear allows 60 POST requests per 10-second window. Sume budgets writes and reads per minute by plan: 120 to 1200 writes, reads 40 times higher.
- Bannerbear video tools mapped to Sume endpoints, gap by gap
Bannerbear lists eleven video tools. Sume covers trim, crop, concat, captions and color via filters; picture-in-picture and GIF previews are not documented.
- Beatoven maestro music and SFX API vs Sume Music Router
Beatoven's API makes music and sound effects from text. Sume's Music Router makes one track per call at a fixed $0.125 and has no SFX route. The differences.
- Best TTS model right now: the leaderboard versus Sume's router
Eleven v4 leads the Artificial Analysis TTS board today. What that means if you generate speech through Sume, whose router serves Sonic models only.
Written by Sume