Voice clone consent record: what to keep before you upload audio
Microsoft says MAI-Voice-2.1 matches a voice from a few seconds of audio. An eight-field consent record, a checker that blocks gaps, and what Sume's docs cover.

Before you upload anyone's reference audio to a voice-matching model, keep a consent record with eight fields: who the speaker is, what they agreed to, which product and markets, the term, how to revoke, who holds the recording, the date, and the hash of the audio you used. The shorter the reference clip a model needs, the less friction there is, and the more the paperwork matters. Microsoft's MAI-Voice-2.1 page lists instant voice matching from short reference clips with no fine-tuning, and its launch post says a voice can be cloned from just a few seconds of reference audio (read 2026-10-03).
This post is a record-keeping pattern, not legal advice. Rules on voice likeness differ by place and change, so have counsel review the wording you actually use.
What the vendor pages do and do not say
The pages I read for MAI-Voice-2.1 describe the capability, not a consent mechanism. The model page names zero-shot voice prompting and granular emotion control. An earlier Microsoft announcement for its first MAI voice model says customers can create their own custom voice in Foundry with a few seconds of audio and that the models were red-teamed and ship with guardrails and governance controls (Microsoft AI, read 2026-10-03). It does not spell out a consent workflow in the text I could read. Before you rely on a vendor's controls, read its current terms for the exact consent obligations and where they sit between you and the vendor.
On Sume's side, the router docs mention pro voice clones only to say that sonic-preview rejects them with voice_model_mismatch. The docs I read for the TTS Router and the other audio surfaces do not describe a consent flow or a verification step for voice creation. Treat consent as something you collect and store, not something the API enforces for you.
The eight fields
The hash matters more than it looks. If a speaker later says a clip was used without agreement, a hash ties the clip you hold to the clip you uploaded, and a file name does not.
| Field | Example | Why you keep it |
|---|---|---|
| Speaker | Full name and contact | Identify who agreed |
| Scope | Product narration in named languages | Limits what you may synthesize |
| Markets and channels | Ads, social, support line | Different rules and risk per channel |
| Term | 12 months from signing | Needs an end date |
| Revocation | Email to a named address, honored in 10 days | Gives the speaker an exit |
| Reference audio hash | SHA-256 of the exact file | Proves which recording was used |
| Recording holder | Who stores the original | Deletion has an owner |
| Signed on | Date and signer | Dates the agreement |
A gate that refuses to proceed
This is a small checker you can put in front of any upload step. It fails when a required field is empty, when the term has already ended, or when the audio file does not match the recorded hash.
import hashlib
import json
from datetime import date
REQUIRED = ["speaker", "scope", "markets", "term_ends", "revocation",
"audio_sha256", "holder", "signed_on"]
def check(record, audio_bytes, today=None):
today = today or date.today()
problems = [f"missing {k}" for k in REQUIRED if not record.get(k)]
if record.get("term_ends") and date.fromisoformat(record["term_ends"]) < today:
problems.append("consent term has ended")
digest = hashlib.sha256(audio_bytes).hexdigest()
if record.get("audio_sha256") and record["audio_sha256"] != digest:
problems.append("audio does not match recorded hash")
return problems
audio = b"example reference audio bytes"
record = {
"speaker": "Jordan Example",
"scope": "Product narration, English and Spanish",
"markets": "Paid social, product pages",
"term_ends": "2027-09-30",
"revocation": "Email consent@example.com; honored within 10 days",
"audio_sha256": hashlib.sha256(audio).hexdigest(),
"holder": "Brand legal",
"signed_on": "2026-10-03",
}
print(json.dumps(check(record, audio, date(2026, 10, 3))))
print(json.dumps(check({**record, "scope": ""}, audio, date(2026, 10, 3))))
Questions a good record can answer
Test the record with four questions a lawyer, a platform reviewer or the speaker themselves might ask. Can you show, in under a minute, who agreed and to what? Can you show which audio file was used and that it is the one the speaker approved? Can the speaker end the arrangement, and does the record say how long you take to act? And if the voice is used somewhere the scope did not list, would anyone notice before it ships?
The last one is a process question more than a paperwork one. If one person can generate with a voice and publish without a second look, the scope field is decoration. Pair the record with a review step for any use that is not on the list, and keep the reviewer's name in the same place you keep the job id.
Short reference clips also raise a quality question. A few seconds of audio carries room noise, a mood and a microphone. Record a clean sample in the speaker's own words, say what it is for, and keep that original file under the retention rule you wrote down.
Where to look in your pipeline
Put the check where the audio enters, not where it leaves. The cheapest control is that nothing reaches a voice-creation step without a passing record, because after that the voice exists. Keep the record next to every job that used the voice, which the next steps in your pipeline can do by storing the job id and the record id together.
Sources
Related posts
More in Use cases
- Voiceover-only Short: silent B-roll plus a TTS spine in Timeline
A Short with no on-camera speech is narration over B-roll. In Timeline the narration is the audio spine, the clips are slots, and audio sets the length.
- Walmart Recognized Reviewer: under 15 reviews, 70% content score
Walmart's Recognized Reviewer now covers items under 15 reviews, if the content quality score is 70% or higher. What to fix first, and what Sume can make.
- Weekly AI release-notes video: 52 x 30 seconds, annual cost on Sume
A weekly 30-second AI presenter release-notes video for a year is 1,560 seconds: $382.20 on Plus, $287.04 on Standard, with a $6.50 music bed and $0.95 avatar.
- A weekly host video while Tavus Griffin is preview-only: Sume, priced
Tavus describes Griffin as a preview with no date or price. For a weekly on-camera update today, Sume Avatar 1.0 costs $3.68 to $11.00 for 20 seconds.
Written by Sume