MAI-Voice-2.1 is not on Sume: make talking clips with Sume TTS

Microsoft MAI-Voice-2.1 launched 2026-10-01 but Sume does not list it. Here is the path that ships: Sume TTS audio, then H3 Max lip sync on a still.

5 min readSume
All posts

Sume does not list Microsoft's MAI-Voice-2.1, so you cannot call it through Sume. To make a talking clip on Sume, generate the voice line with Sume's own text-to-speech, then submit the audio and a still to the H3 Max lip-sync route, or let the Avatar 1.0 talking-video route do both from a script. A tracker records MAI-Voice-2.1 as released on 2026-10-01 with 23 languages across 26 locales, one voice across languages and $22 per million characters, plus a Flash variant with a vendor claim of 150 ms end to end at $15 per million (Digital Applied tracker, read 2026-10-06).

Those are Microsoft's figures as reported by the tracker, and they are for audio only. This post does not compare them with Sume's TTS price; the Sume TTS catalog at GET /v1/tts-router/models shows live per-character rates.

Why does the audio source matter for lip sync?

Sume's lip-sync route accepts only audio that is hosted on Sume: the audio_url goes through a Sume-host and 10 MB preflight, and the clip must run 5 to 14.8 seconds. A file made elsewhere is not accepted by its own URL. The docs for other Sume audio routes say to import files first with POST /v1/media-imports, which gives you a media.sume.com URL, but the lip-sync page does not say that an imported file is accepted, so test one before you plan around it. Either way, a voice from another vendor is an extra step and not a drop-in for this route.

The supported order is the one Sume's own guidance gives: TTS first, then a lip-sync model, never a video model with narration laid under it. Sume's TTS job returns a durable media.sume.com audio artifact, and that URL is what you pass as audio_url. Measure the audio length from the TTS result and send it as duration_seconds; the API rejects a value outside 5 to 14.8 rather than clamping it.

What is the two-call path?

The Sume-only path from script to talking clip. Routes from docs.sume.com Models and Generate avatar video, read 2026-10-06.
StepCallOutput
1. Make the voice linePOST /v1/tts-1.0/generate with a transcript and a voice or avatarAudio artifact on media.sume.com
2. Read its lengthGET /v1/jobs/:id/resultMeasured seconds
3. Animate the stillPOST /v1/minimax/h3-max/lip-sync with image_url or avatar_handle, audio_url, duration_secondsTalking clip
Shortcut for a whole scriptPOST /v1/avatar-1.0/talking-video with avatar_handle and scriptOne clip, 4 to 60 s

When is the shortcut better?

Use the two-call path when you want to control the voice, check the audio before paying for video, or reuse one audio file across stills. Use the talking-video route when you want the avatar's own voice and a single job. If the avatar has no usable voice, the API answers with a 400 and the avatar voice post lists fixes.

Languages are the one place MAI's pitch overlaps a decision: if you need a voice that carries across 23 languages, check what languages Sume's catalog lists for your target, and run a sample. This post makes no claim about parity. If your line is longer than 14.8 seconds, split it at sentence boundaries and submit one job per part, each at least 5 seconds. See the window post for the cut.

How do you compare voices fairly?

If you are choosing between a trending voice model and Sume's TTS, run one script through each and compare the audio blind. Judge pronunciation of your product names, pacing and how the sentence ends, since lip sync will expose awkward endings. Only the Sume audio can go straight into the lip-sync route, so factor in the extra import step and test that a different voice would need.

Keep price comparisons honest. The tracker's $22 and $15 per million characters are per character of text, while Sume's lip-sync is billed per second of video. They are not interchangeable units, and a voice line is only one part of a finished talking clip.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume