MAI-Voice-2.1 clones from seconds of audio: what Sume's API takes

Microsoft says MAI-Voice-2.1 clones a voice from seconds of audio. Sume does not carry it, and voice cloning on Sume is app-only; the API takes voice ids.

5 min readSume
All posts

Sume does not list MAI-Voice, and you cannot clone a voice through the Sume API at all: cloning happens in the Sume app, and the API only accepts ids of voices that already exist. Microsoft AI's October 1, 2026 post says MAI-Voice-2.1 supports voice cloning from seconds of reference audio.

Sources: Microsoft AI's announcement (read 2026-10-10) and Sume's API reference.

What does Microsoft say?

The post introduces three models. MAI-Voice-2.1 is described as covering 23 languages and 26 locales, with one voice across languages and cloning from seconds of audio, at $22 per million characters. MAI-Voice-2.1-Flash is the faster sibling at $15 per million characters. The post also lists availability on OpenRouter, Foundry, a MAI Playground and Vercel, with LiveKit listed as coming soon.

MAI voice models as stated by Microsoft AI (read 2026-10-10)
ModelStated detailStated price
MAI-Voice-2.123 languages, 26 locales, cloning from seconds of audio$22 per 1M characters
MAI-Voice-2.1-Flash45 s of audio at 150 ms end to end$15 per 1M characters

What does Sume's TTS take instead?

Sume TTS takes a voice by avatar_id or avatar_handle (the avatar's voice must have voice.status ready) or by voice.id, a TTS voice UUID or a voi_ id from your Voices library. Any other id shape fails with 400 invalid_voice_id before queueing, so no credits are reserved. Cloning a new voice from a recording is done in the Sume app; after that the voice has an id you can use over the API.

  • API: use an existing voice id or ready avatar voice.
  • App: create the voice from your recording.
  • Language: set language for non-English transcripts.

How do the prices compare?

Sume bills $0.0475 per 1,000 characters, which is $47.50 per million, rounded up to a whole cent per job. Microsoft's stated $22 per million is a different shape: it is a per-million vendor rate, and Sume's figure includes the platform's margin and the job minimum. A fair read is that the two are not like for like, because Sume also hosts the result and feeds it into captions, timeline and lip-sync jobs.

Rates side by side (Microsoft AI read 2026-10-10; Sume catalog checked 2026-10-10)
ItemPer 1M charactersFor a 1,000-character script
MAI-Voice-2.1 (stated)$22$0.022
MAI-Voice-2.1-Flash (stated)$15$0.015
Sume TTS$47.50$0.0475 rounded up to $0.05

Which should you use?

Pick MAI-Voice if you need its cloning flow and its language spread, and you are happy to call Microsoft's endpoints. Pick Sume when the audio is one step of a video pipeline and you want one hosted result usable by the next job. Consent matters for any cloned voice: only clone a voice you have the right to use.

What is a safe way to use a cloned voice?

Whatever tool makes the clone, the controls are the same. Get written permission from the person whose voice it is, record what the voice may be used for, and keep the reference recording and consent note with the project. On Sume, a voice created in the app becomes an id you can pass to the API, so the consent record should travel with that id.

Because the API cannot create voices, an automated pipeline on Sume cannot accidentally clone someone from a URL. That is a limit, but it also means every voice in your API calls was set up by a person in the app first.

  • Keep the consent note next to the voice id in your config.
  • Do not reuse a voice for a use the person did not agree to.
  • Re-check voice.status before a batch if an avatar voice was edited.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume