Voice clone consent record: what to keep before you upload audio

Microsoft says MAI-Voice-2.1 matches a voice from a few seconds of audio. An eight-field consent record, a checker that blocks gaps, and what Sume's docs cover.

6 min readSume
All posts

Before you upload anyone's reference audio to a voice-matching model, keep a consent record with eight fields: who the speaker is, what they agreed to, which product and markets, the term, how to revoke, who holds the recording, the date, and the hash of the audio you used. The shorter the reference clip a model needs, the less friction there is, and the more the paperwork matters. Microsoft's MAI-Voice-2.1 page lists instant voice matching from short reference clips with no fine-tuning, and its launch post says a voice can be cloned from just a few seconds of reference audio (read 2026-10-03).

This post is a record-keeping pattern, not legal advice. Rules on voice likeness differ by place and change, so have counsel review the wording you actually use.

What the vendor pages do and do not say

The pages I read for MAI-Voice-2.1 describe the capability, not a consent mechanism. The model page names zero-shot voice prompting and granular emotion control. An earlier Microsoft announcement for its first MAI voice model says customers can create their own custom voice in Foundry with a few seconds of audio and that the models were red-teamed and ship with guardrails and governance controls (Microsoft AI, read 2026-10-03). It does not spell out a consent workflow in the text I could read. Before you rely on a vendor's controls, read its current terms for the exact consent obligations and where they sit between you and the vendor.

On Sume's side, the router docs mention pro voice clones only to say that sonic-preview rejects them with voice_model_mismatch. The docs I read for the TTS Router and the other audio surfaces do not describe a consent flow or a verification step for voice creation. Treat consent as something you collect and store, not something the API enforces for you.

The eight fields

The hash matters more than it looks. If a speaker later says a clip was used without agreement, a hash ties the clip you hold to the clip you uploaded, and a file name does not.

A suggested record, not a legal template. Capability facts from Microsoft's MAI-Voice-2.1 page, read 2026-10-03.
FieldExampleWhy you keep it
SpeakerFull name and contactIdentify who agreed
ScopeProduct narration in named languagesLimits what you may synthesize
Markets and channelsAds, social, support lineDifferent rules and risk per channel
Term12 months from signingNeeds an end date
RevocationEmail to a named address, honored in 10 daysGives the speaker an exit
Reference audio hashSHA-256 of the exact fileProves which recording was used
Recording holderWho stores the originalDeletion has an owner
Signed onDate and signerDates the agreement

A gate that refuses to proceed

This is a small checker you can put in front of any upload step. It fails when a required field is empty, when the term has already ended, or when the audio file does not match the recorded hash.

import hashlib
import json
from datetime import date

REQUIRED = ["speaker", "scope", "markets", "term_ends", "revocation",
            "audio_sha256", "holder", "signed_on"]

def check(record, audio_bytes, today=None):
    today = today or date.today()
    problems = [f"missing {k}" for k in REQUIRED if not record.get(k)]
    if record.get("term_ends") and date.fromisoformat(record["term_ends"]) < today:
        problems.append("consent term has ended")
    digest = hashlib.sha256(audio_bytes).hexdigest()
    if record.get("audio_sha256") and record["audio_sha256"] != digest:
        problems.append("audio does not match recorded hash")
    return problems

audio = b"example reference audio bytes"
record = {
    "speaker": "Jordan Example",
    "scope": "Product narration, English and Spanish",
    "markets": "Paid social, product pages",
    "term_ends": "2027-09-30",
    "revocation": "Email consent@example.com; honored within 10 days",
    "audio_sha256": hashlib.sha256(audio).hexdigest(),
    "holder": "Brand legal",
    "signed_on": "2026-10-03",
}
print(json.dumps(check(record, audio, date(2026, 10, 3))))
print(json.dumps(check({**record, "scope": ""}, audio, date(2026, 10, 3))))

Questions a good record can answer

Test the record with four questions a lawyer, a platform reviewer or the speaker themselves might ask. Can you show, in under a minute, who agreed and to what? Can you show which audio file was used and that it is the one the speaker approved? Can the speaker end the arrangement, and does the record say how long you take to act? And if the voice is used somewhere the scope did not list, would anyone notice before it ships?

The last one is a process question more than a paperwork one. If one person can generate with a voice and publish without a second look, the scope field is decoration. Pair the record with a review step for any use that is not on the list, and keep the reviewer's name in the same place you keep the job id.

Short reference clips also raise a quality question. A few seconds of audio carries room noise, a mood and a microphone. Record a clean sample in the speaker's own words, say what it is for, and keep that original file under the retention rule you wrote down.

Where to look in your pipeline

Put the check where the audio enters, not where it leaves. The cheapest control is that nothing reaches a voice-creation step without a passing record, because after that the voice exists. Keep the record next to every job that used the voice, which the next steps in your pipeline can do by storing the job id and the record id together.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume