MAI-Voice-2.1 clones from seconds: how Sume clones from a clip
Microsoft says MAI-Voice-2.1 clones from a few seconds of audio with consent guardrails. In Sume Agents, voices_create clones from an https audio clip.

Microsoft's October 1 post says MAI-Voice-2.1 supports voice cloning using just a few seconds of reference audio, and that it has built-in consent guardrails that prevent misuse. Sume does not run MAI-Voice-2.1. Its clone path is the voices_create tool in Sume Agents, which makes one workspace voice from an https audio clip, and the public TTS routes then speak with that voice.
How a Sume clone is made
Give voices_create an audio_url and it creates exactly one voice. The clip option cannot be mixed with a persona, an invent request, or variants, and a clone needs no persona text. name is optional and gender defaults to female. Set language to the language of the clip. The row lands in Assets, Voices, with source_kind clone.
From video to voice id
If the source is a video, run audio_detach first and pass its audio_url to voices_create. The create call is paid, so confirm spend when it matters. The tool waits until the row is ready and returns a cartesia_voice_id. Then tts_create speaks with that value as voice.id.
What to check in the REST docs
The public REST TTS routes select an existing voice by avatar or voice.id. This page describes the Agents tool path, so do not look for a clone endpoint in the REST reference.
Both paths, read 2026-10-06:
| Item | MAI-Voice-2.1 (vendor) | Sume Agents (code) |
|---|---|---|
| Reference input | A few seconds of audio | An https audio clip |
| Consent | Built-in guardrails | Set by you and the provider's terms of use |
| Output | Voice usable in 23 languages | One workspace voice with a language tag |
| Price | $22 per 1M characters for speech | Per-character TTS after a paid create |
Before you clone anything
Get the speaker's written agreement, and keep it with the project. Tell them where the voice will be used. A clone is a copy of a person, and a platform's rules on synthetic voices are separate from your own.
Use a clean clip. Remove music and other voices first. If the clip is a video, audio_detach gives you a wav. Run a short test line in the language you want before you spend on a long script, and listen for the accent.
Finally, remember that a voice you made in one workspace is a workspace asset. It can be mentioned later in Agents and selected by id for speech, so decide who in your team may use it, and delete it when the consent ends.
Use only voices you have the right to use. Whatever model you pick, keep the written consent of the speaker, and disclose synthetic voices where a platform or law asks for it.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 on Foundry and Vercel; Sume's TTS Router is Sonic only
Microsoft lists Foundry, MAI Playground, Vercel and OpenRouter for MAI-Voice-2.1. Sume's TTS Router lists only Cartesia Sonic ids, so check the catalog.
- MAI-Voice-2.1 or Sume TTS? Three questions that decide it
Live speech, extra outputs or lowest price per character? A short guide to MAI-Voice-2.1 and Flash against the Sume TTS Router, with 10M-character math.
- Shortest AI video clip by model: MiniMax H3 4 s, Sume row 5 s
MiniMax lists 4 to 15 seconds for H3. Sume's row starts at 5. Here are the minimum durations in the catalog and how to trim a longer clip.
- MiniMax H3 open weights vs a hosted lip-sync API: what you take on
MiniMax released H3 open weights on 2026-08-03. Self-hosting is not the same as a hosted still-plus-audio lip-sync route. What each choice makes you own.
Written by Sume