Licensed data, consented voice talent: 6 questions for your TTS vendor
Decagon says Chord trained on licensed data and consented voice talent. Six questions to ask any text-to-speech vendor before you ship an AI voice.

Ask your TTS vendor six things before you ship a synthetic voice: what the training data was, whether the voices are consented, how cloning is gated, whether the audio is watermarked, what the commercial terms are, and which jobs the voice was tuned for. Three vendors made public statements on this in the last two weeks, and they differ.
What each vendor says on its own pages
The statements below come from each vendor's own page, read on 2026-10-07.
| Vendor | Statement |
|---|---|
| Decagon (Chord) | Trained on licensed data and consented voice talent |
| Microsoft (MAI-Voice-2.1) | Built-in guardrails so only authorized, consented voices can be used; cloning is gated access |
| Google (Gemini 3.8 TTS) | Voice replication from a 30-second sample with consent verification; audio watermarked with SynthID |
| Mistral (Voxtral TTS) | Weights under CC BY-NC 4.0; API available through Mistral Studio |
The six questions
Use these as a checklist in a vendor email.
- What data trained the voice, and was the voice talent paid and consenting?
- Can a customer clone a voice, and what proof of consent does the vendor require?
- Is the output watermarked, and does the mark survive trimming and re-encoding?
- What are the commercial terms for ads, and is the license for the weights or the API?
- Was the voice tuned for your use, such as narration, or for another, such as support calls?
- Which terms apply if you modify the voice or edit the output?
Why the last two matter most
Decagon describes Chord as post-trained on real-world customer-experience conversations, which points to support calls. Microsoft says MAI-Voice-2.1 prioritizes expression over latency, and its Flash variant is aimed at agents. A vendor's tuning tells you where it is likely to shine and where it will not.
The license matters too: Mistral lists Voxtral TTS weights as CC BY-NC 4.0, which is non-commercial, while the paid API is a separate route.
Where Sume fits
Sume's TTS surface takes a voice selector from the workspace's voice library or an avatar, and it records what it used. A completed text_to_speech job stores the model id, voice, language, output format and the synthesis settings, so you can show which voice made which file. For consent and licensing questions about a specific voice, ask Sume support before you rely on it in a campaign; this post only reports what the job record shows.
Sources
Related posts
More in Comparisons
- Veo 3.1 price per second vs Sume's Video Router models
Google lists Veo 3.1 at $0.05 to $0.60 per second. Sume does not offer Veo 3.1 in its v1 catalog; here is the 8-second, 100-clip math for what it does.
- Veo 3.1 Standard, Fast, Lite: price per 8-second clip vs Sume
An 8-second Veo 3.1 clip costs $3.20 Standard, $0.80 Fast and $0.40 Lite at 720p on Google's page. Wan 3.0 and Omni Flash on Sume are $1.25 before the fee.
- A 10-second 9:16 ad: what each Sume video model costs
Ten seconds of 9:16 video on Sume: H3 768p $0.75, H3 Max 768p $1.00, Wan 720p $1.25, Omni 720p $1.25. All except Wan carry always-on audio.
- Voice agent, TTS API or rendered narration? Pick by output
Decagon Voice 3, MAI-Voice-2.1 Flash, Gemini 3.8 TTS and Sume jobs solve different problems. A decision table for choosing by output, October 2026.
Written by Sume