AI voice disclosure: OpenAI rule, cloning consent, Voxtral license
What the OpenAI, Microsoft and Mistral pages say about disclosing AI speech, cloning a voice and non-commercial weights, and a publish-time checklist.

Before you publish a narrated video, check three things the vendors say in public: whether you must tell listeners the voice is synthetic, whether cloning a real person needs their consent, and whether the model's license allows your use. OpenAI's guide says its usage policies require a clear disclosure that the voice is AI-generated. Microsoft describes MAI-Voice-2.1 as cloning a voice from a short reference clip, and the page I read states no consent terms, so you supply them yourself. Mistral's open weights are CC BY-NC 4.0. This is a read of vendor pages, not legal advice.
What each page says
These are the vendor statements as read on 2026-10-04.
| Vendor | What the page says | What it means for you |
|---|---|---|
| OpenAI TTS guide | Usage policies require a clear disclosure to end users that the TTS voice is AI-generated, not a human voice | Add a disclosure where listeners will see it |
| Microsoft MAI-Voice-2.1 | Voice capture from a short reference clip; the page I read gives no consent terms | Get and keep the speaker's written consent before you clone, whatever the vendor page says |
| Mistral Voxtral TTS | Open weights under CC BY-NC 4.0; API at $0.016 per 1,000 characters | Weights are for non-commercial use; paid work needs the API or a separate agreement |
A publish-time checklist
Run it per video, because the platform you publish on can have its own rule on top of the vendor's.
- Name the voice vendor and model in your project notes, with the date you generated it.
- If the vendor requires disclosure, put it in the description or on screen, not only in a terms page.
- If you cloned a person, keep a written consent from that person for the specific use.
- If you ran open weights, confirm the license allows paid work; CC BY-NC does not.
- Check the destination platform's label settings for synthetic audio before upload.
What Sume does and does not decide for you
Sume generates the speech and returns a hosted artifact; the policy about telling viewers is yours and the model vendor's. The basics page describes the workflow, and the jobs and results page shows how to read a finished job. Neither page makes a disclosure decision for you.
If you want the disclosure on screen, one practical route is to burn it as text with the video captions endpoint, which accepts authored cues with text, start and end. That endpoint charges $0.20 per accepted job for videos up to 60 seconds under the current fixed estimate, so confirm the live price in the catalog. A cue such as a short line on the first two seconds is enough for the mechanics; the wording and placement should follow the rule you are meeting.
Sources
Related posts
More in Use cases
- AI wall art canvas: gpt-image-2.5 pixels per inch, 8x10 to 24x36
The biggest 2:3 image gpt-image-2.5 can send is 2336x3504, which is 97 pixels per inch at 24x36. Here is the dpi for five canvas sizes and when to upscale.
- AI wine label design via API: artwork from the model, text from code
For a wine label, generate only the artwork with gpt-image-2.5, ask for a clear panel, and set the name, vintage and regulatory lines in code so they are exact.
- Four shoppable videos per ASIN: demo, how-to, lifestyle, unboxing
Amazon allows up to 4 shoppable videos per ASIN for sellers with a professional account and three months of history. A four-clip plan batched on Sume.
- Amazon allows four shoppable videos per ASIN: eligibility and plan
Amazon allows up to four shoppable videos per ASIN for professional sellers with three months of history. A demo, how-to, lifestyle and unboxing plan.
Written by Sume