Use a MAI-Voice-2.1 clip as avatar audio? Sume needs its own file

A talking still on Sume takes Sume-hosted audio under 10 MB. A MAI-Voice-2.1 clip sits elsewhere, so make the line with Sume TTS instead. Sizes and limits.

5 min readSume
All posts

Not directly. Sume's talking-still models read audio that is hosted on Sume and smaller than 10 MB, and a clip made by MAI-Voice-2.1 lives in Microsoft's service or in your own storage. To get a lip-synced presenter on Sume, synthesize the line with Sume TTS and pass its audio_url to the avatar job.

Microsoft's 2026-10-01 launch lists MAI-Voice-2.1 at $22 per 1M characters, and MAI-Voice-2.1-Flash at $15 per 1M characters with a 45-second audio ceiling per call (Microsoft AI, read 2026-10-05). Neither price changes what the Sume avatar routes accept.

What each Sume talking-still route asks for

The three talking-still routes in Sume's contract each state their audio rule in the catalog text. Duration limits differ, and so does the hosting rule.

Audio rules by Sume route (read 2026-10-05)
RouteAudio lengthAudio source
VEED Fabric 1.0 (veed/fabric-1.0)Up to 300 secondsSume-hosted, under 10 MB
Sume Avatar 1.0 Fabric (experimental)4 to 15 seconds of outputSume-hosted, under 10 MB
MiniMax H3 Max Lip Sync5 to 14.8 secondsSume-hosted, under 10 MB

The call shape

The Models page describes the Fabric call: send audio_url, a measured duration_seconds, and one visual source. It also states that video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is a talking-still job with a TTS file, not a video clip with narration under it.

How many seconds fit in 10 MB

The size cap becomes a length cap once you pick a format. Taking 10 MB as 10,000,000 bytes, a mono 44.1 kHz 16-bit WAV runs 88,200 bytes per second, so about 113 seconds fit. A 128 kbps MP3 runs 16,000 bytes per second, so about 625 seconds fit, which is more than the 300-second Fabric ceiling. That is simple arithmetic, not a measured Sume limit, so check your real file.

Seconds that fit under 10,000,000 bytes (read 2026-10-05)
FormatBytes per secondSeconds in 10 MB
WAV, mono, 44.1 kHz, 16-bit88,200113
WAV, stereo, 44.1 kHz, 16-bit176,40056
MP3, 128 kbps16,000625

The Sume path

Sume TTS defaults to MP3 at 44.1 kHz and 128 kbps, so a short line is well under the cap. Pick the voice with avatar_id or avatar_handle and Sume resolves that avatar's ready TTS voice. A Sume TTS request costs $0.0475 per 1,000 characters at list, so a 600-character line is $0.0285.

If you need the avatar to speak in a MAI voice, there is no route for it. Keep the MAI audio for other uses, for example a voice-only deliverable, and ask the avatar to speak the same text in a Sume voice.

Check before you render

Before you spend on a render, measure the clip you plan to send. The route's catalog text says the audio must be under 10 MB, so convert to MP3 at 128 kbps if a WAV is too big.

Keep the line under the 45-second Flash ceiling only if you plan to compare against MAI Flash. On Sume the avatar route, not the TTS model, sets the length limit.

There is also a rights and consent angle. A cloned or reference-matched voice from another service comes with that service's terms, and Microsoft's launch note mentions consent guardrails for cloning across languages. A presenter voice on Sume is tied to a workspace avatar whose voice status is ready, which keeps the voice, the face and the script in one place you can audit later. Whichever route you pick, write down which service made the audio before the clip ships.

For a quick decision: if the viewer must see a face speaking, make the audio on Sume. If the viewer only hears a voice, a MAI file is fine, and Sume STT or the caption routes can still read it from a public HTTPS URL.

Why this costs little

Cost sits on both sides of the choice. A Sume TTS line of 600 characters is $0.0285 at the list rate, and the talking-still render is billed on top, per second of audio. Reusing a Microsoft file would save the 2.85 cents but not the render. The saving from bringing your own audio is small, and a mismatched pipeline can cost more in rework than that.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume