Licensed voices only: what an ad team checks before take one
Microsoft says MAI-Voice-2 synthesizes only authorized, licensed voices. Here is the rights checklist to run for any AI voiceover, and what Sume's docs cover.

Before the first take of an AI voiceover, confirm three things in writing: whose voice it is, what that person agreed to, and where the voice id came from. Microsoft's MAI-Voice-2 announcement says only authorized, licensed voices can be synthesized in production and that no unlicensed voice cloning is possible (read 2026-10-05), which is the right bar for any engine.
What Microsoft states
The MAI-Voice-2 announcement describes zero-shot voice prompting from 5 to 60 seconds of reference audio, with consent enforced at the system level (read 2026-10-05). The MAI-Voice-2.1 model page lists instant voice matching from a short reference clip with no fine-tuning, on both the Standard and Flash variants (read 2026-10-05). That page does not spell out licensing terms, so a team should not assume what is not on it.
What Sume's docs cover
Sume's TTS docs describe how a voice is selected, not a consent workflow. You choose a voice with voice.id (a TTS voice UUID) or through an avatar id or handle, and the router does not create a second voice namespace. Sume's TTS Router doc also notes that the beta sonic-preview id rejects pro voice clones with voice_model_mismatch, which is a reason to keep reads for a cloned spokesperson on a stable id such as sonic-3.6.
Because the docs do not describe consent capture, keep your own record. The checklist below is for the team, not a Sume feature.
The pre-take checklist
Run it once per voice, not once per ad.
| Check | What to keep | Why |
|---|---|---|
| Whose voice | Name, role, a dated sign-off | A voice id alone does not prove who agreed |
| Scope of consent | Uses, languages, media, term | A consent for English radio is not consent for Spanish paid social |
| Source of the id | Library voice, avatar voice or a clone, and the date made | You need to find every read if consent is withdrawn |
| Engine pin | job.model from the job record | Router jobs record the requested id, so a clone's reads stay traceable |
| Language | Voice language against request language | Sume returns 409 tts_voice_language_mismatch before a charge, unless you confirm |
| Disclosure | Where the ad says the voice is synthetic | Platform and local rules differ, so check yours |
Why the language check belongs here
A spokesperson who agreed to English reads has not agreed to a Spanish read in their voice. Sume's TTS docs add a double-check: a known mismatch between the voice's primary language and the request language returns HTTP 409 tts_voice_language_mismatch before any job or charge, and you retry with confirm_language_mismatch: true only after asking. Use the 409 as a trigger to check scope, not a prompt to click through.
Regional tags compare by primary language, so es-MX against a Spain Spanish voice passes without a warning. If regional consent matters, your records must say so, because the API will not.
What to ask a vendor, any vendor
Microsoft's page answers the first question in its own words. The others are not on the pages read here, so they belong in your contract, not your assumptions.
- Can an unauthorized voice be synthesized, and how is that enforced?
- Who holds the licensing rights to the output, and are they transferable to an advertiser?
- Is there a published retention period for reference audio?
- What happens to existing reads when a voice is retired?
When consent is withdrawn
Plan the takedown on day one. Tag each job with metadata (stored with the request, never sent to the provider) holding the voice id and the campaign id, so a search by voice returns every read. Re-rendering a 1,350-character read costs 7 cents, so replacing 100 reads costs $7 in TTS, and the longer cost is re-rendering the videos around them.
Sources
Related posts
More in Use cases
- LinkedIn Lead Gen Form hidden fields: tag each lead with its variant
A LinkedIn Lead Gen Form allows up to 20 hidden fields. Keep your own index-to-variant table from the Sume bulk queue so every lead traces to one creative.
- LinkedIn Lead Gen Form image 552x200: make it from a 3:1 render
LinkedIn recommends a 552px by 200px form image. Sume has no 552:200 ratio, so ask Ideogram for 3:1, center-crop about 8% of the width, and resize.
- LinkedIn video ad intro text 150 characters, headline 70: test copy
LinkedIn video ads show 150 characters of intro text and a 70-character headline. Test copy on one rendered clip and save Sume jobs for hook variants.
- Two listing photos, one room transition: first and last frame
Use two listing photos as first and last frame of a 6-second Gemini Omni Flash 1.1 clip on Sume: $0.75 at 720p, $1.125 at 1080p. Call and checks.
Written by Sume