AI voice disclosure: OpenAI rule, cloning consent, Voxtral license

What the OpenAI, Microsoft and Mistral pages say about disclosing AI speech, cloning a voice and non-commercial weights, and a publish-time checklist.

5 min readSume
All posts

Before you publish a narrated video, check three things the vendors say in public: whether you must tell listeners the voice is synthetic, whether cloning a real person needs their consent, and whether the model's license allows your use. OpenAI's guide says its usage policies require a clear disclosure that the voice is AI-generated. Microsoft describes MAI-Voice-2.1 as cloning a voice from a short reference clip, and the page I read states no consent terms, so you supply them yourself. Mistral's open weights are CC BY-NC 4.0. This is a read of vendor pages, not legal advice.

What each page says

These are the vendor statements as read on 2026-10-04.

Disclosure, consent and license statements by vendor page (read 2026-10-04)
VendorWhat the page saysWhat it means for you
OpenAI TTS guideUsage policies require a clear disclosure to end users that the TTS voice is AI-generated, not a human voiceAdd a disclosure where listeners will see it
Microsoft MAI-Voice-2.1Voice capture from a short reference clip; the page I read gives no consent termsGet and keep the speaker's written consent before you clone, whatever the vendor page says
Mistral Voxtral TTSOpen weights under CC BY-NC 4.0; API at $0.016 per 1,000 charactersWeights are for non-commercial use; paid work needs the API or a separate agreement

A publish-time checklist

Run it per video, because the platform you publish on can have its own rule on top of the vendor's.

  • Name the voice vendor and model in your project notes, with the date you generated it.
  • If the vendor requires disclosure, put it in the description or on screen, not only in a terms page.
  • If you cloned a person, keep a written consent from that person for the specific use.
  • If you ran open weights, confirm the license allows paid work; CC BY-NC does not.
  • Check the destination platform's label settings for synthetic audio before upload.

What Sume does and does not decide for you

Sume generates the speech and returns a hosted artifact; the policy about telling viewers is yours and the model vendor's. The basics page describes the workflow, and the jobs and results page shows how to read a finished job. Neither page makes a disclosure decision for you.

If you want the disclosure on screen, one practical route is to burn it as text with the video captions endpoint, which accepts authored cues with text, start and end. That endpoint charges $0.20 per accepted job for videos up to 60 seconds under the current fixed estimate, so confirm the live price in the catalog. A cue such as a short line on the first two seconds is enough for the mechanics; the wording and placement should follow the rule you are meeting.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume