MAI-Voice-2.1 clones from seconds: how Sume clones from a clip

Microsoft says MAI-Voice-2.1 clones from a few seconds of audio with consent guardrails. In Sume Agents, voices_create clones from an https audio clip.

4 min readSume
All posts

Microsoft's October 1 post says MAI-Voice-2.1 supports voice cloning using just a few seconds of reference audio, and that it has built-in consent guardrails that prevent misuse. Sume does not run MAI-Voice-2.1. Its clone path is the voices_create tool in Sume Agents, which makes one workspace voice from an https audio clip, and the public TTS routes then speak with that voice.

How a Sume clone is made

Give voices_create an audio_url and it creates exactly one voice. The clip option cannot be mixed with a persona, an invent request, or variants, and a clone needs no persona text. name is optional and gender defaults to female. Set language to the language of the clip. The row lands in Assets, Voices, with source_kind clone.

From video to voice id

If the source is a video, run audio_detach first and pass its audio_url to voices_create. The create call is paid, so confirm spend when it matters. The tool waits until the row is ready and returns a cartesia_voice_id. Then tts_create speaks with that value as voice.id.

What to check in the REST docs

The public REST TTS routes select an existing voice by avatar or voice.id. This page describes the Agents tool path, so do not look for a clone endpoint in the REST reference.

Both paths, read 2026-10-06:

Voice cloning, MAI-Voice-2.1 versus Sume, read 2026-10-06
ItemMAI-Voice-2.1 (vendor)Sume Agents (code)
Reference inputA few seconds of audioAn https audio clip
ConsentBuilt-in guardrailsSet by you and the provider's terms of use
OutputVoice usable in 23 languagesOne workspace voice with a language tag
Price$22 per 1M characters for speechPer-character TTS after a paid create

Before you clone anything

Get the speaker's written agreement, and keep it with the project. Tell them where the voice will be used. A clone is a copy of a person, and a platform's rules on synthetic voices are separate from your own.

Use a clean clip. Remove music and other voices first. If the clip is a video, audio_detach gives you a wav. Run a short test line in the language you want before you spend on a long script, and listen for the accent.

Finally, remember that a voice you made in one workspace is a workspace asset. It can be mentioned later in Agents and selected by id for speech, so decide who in your team may use it, and delete it when the consent ends.

Use only voices you have the right to use. Whatever model you pick, keep the written consent of the speaker, and disclose synthetic voices where a platform or law asks for it.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume