Licensed voices only: what an ad team checks before take one

Microsoft says MAI-Voice-2 synthesizes only authorized, licensed voices. Here is the rights checklist to run for any AI voiceover, and what Sume's docs cover.

5 min readSume
All posts

Before the first take of an AI voiceover, confirm three things in writing: whose voice it is, what that person agreed to, and where the voice id came from. Microsoft's MAI-Voice-2 announcement says only authorized, licensed voices can be synthesized in production and that no unlicensed voice cloning is possible (read 2026-10-05), which is the right bar for any engine.

What Microsoft states

The MAI-Voice-2 announcement describes zero-shot voice prompting from 5 to 60 seconds of reference audio, with consent enforced at the system level (read 2026-10-05). The MAI-Voice-2.1 model page lists instant voice matching from a short reference clip with no fine-tuning, on both the Standard and Flash variants (read 2026-10-05). That page does not spell out licensing terms, so a team should not assume what is not on it.

What Sume's docs cover

Sume's TTS docs describe how a voice is selected, not a consent workflow. You choose a voice with voice.id (a TTS voice UUID) or through an avatar id or handle, and the router does not create a second voice namespace. Sume's TTS Router doc also notes that the beta sonic-preview id rejects pro voice clones with voice_model_mismatch, which is a reason to keep reads for a cloned spokesperson on a stable id such as sonic-3.6.

Because the docs do not describe consent capture, keep your own record. The checklist below is for the team, not a Sume feature.

The pre-take checklist

Run it once per voice, not once per ad.

Rights checklist before an AI voiceover ships (Microsoft pages and Sume TTS docs, read 2026-10-05)
CheckWhat to keepWhy
Whose voiceName, role, a dated sign-offA voice id alone does not prove who agreed
Scope of consentUses, languages, media, termA consent for English radio is not consent for Spanish paid social
Source of the idLibrary voice, avatar voice or a clone, and the date madeYou need to find every read if consent is withdrawn
Engine pinjob.model from the job recordRouter jobs record the requested id, so a clone's reads stay traceable
LanguageVoice language against request languageSume returns 409 tts_voice_language_mismatch before a charge, unless you confirm
DisclosureWhere the ad says the voice is syntheticPlatform and local rules differ, so check yours

Why the language check belongs here

A spokesperson who agreed to English reads has not agreed to a Spanish read in their voice. Sume's TTS docs add a double-check: a known mismatch between the voice's primary language and the request language returns HTTP 409 tts_voice_language_mismatch before any job or charge, and you retry with confirm_language_mismatch: true only after asking. Use the 409 as a trigger to check scope, not a prompt to click through.

Regional tags compare by primary language, so es-MX against a Spain Spanish voice passes without a warning. If regional consent matters, your records must say so, because the API will not.

What to ask a vendor, any vendor

Microsoft's page answers the first question in its own words. The others are not on the pages read here, so they belong in your contract, not your assumptions.

  • Can an unauthorized voice be synthesized, and how is that enforced?
  • Who holds the licensing rights to the output, and are they transferable to an advertiser?
  • Is there a published retention period for reference audio?
  • What happens to existing reads when a voice is retired?

When consent is withdrawn

Plan the takedown on day one. Tag each job with metadata (stored with the request, never sent to the provider) holding the voice id and the campaign id, so a search by voice returns every read. Re-rendering a 1,350-character read costs 7 cents, so replacing 100 reads costs $7 in TTS, and the longer cost is re-rendering the videos around them.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume