Use a MAI-Voice-2.1 clip as avatar audio? Sume needs its own file
A talking still on Sume takes Sume-hosted audio under 10 MB. A MAI-Voice-2.1 clip sits elsewhere, so make the line with Sume TTS instead. Sizes and limits.
Not directly. Sume's talking-still models read audio that is hosted on Sume and smaller than 10 MB, and a clip made by MAI-Voice-2.1 lives in Microsoft's service or in your own storage. To get a lip-synced presenter on Sume, synthesize the line with Sume TTS and pass its audio_url to the avatar job.
Microsoft's 2026-10-01 launch lists MAI-Voice-2.1 at $22 per 1M characters, and MAI-Voice-2.1-Flash at $15 per 1M characters with a 45-second audio ceiling per call (Microsoft AI, read 2026-10-05). Neither price changes what the Sume avatar routes accept.
What each Sume talking-still route asks for
The three talking-still routes in Sume's contract each state their audio rule in the catalog text. Duration limits differ, and so does the hosting rule.
| Route | Audio length | Audio source |
|---|---|---|
VEED Fabric 1.0 (veed/fabric-1.0) | Up to 300 seconds | Sume-hosted, under 10 MB |
| Sume Avatar 1.0 Fabric (experimental) | 4 to 15 seconds of output | Sume-hosted, under 10 MB |
| MiniMax H3 Max Lip Sync | 5 to 14.8 seconds | Sume-hosted, under 10 MB |
The call shape
The Models page describes the Fabric call: send audio_url, a measured duration_seconds, and one visual source. It also states that video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is a talking-still job with a TTS file, not a video clip with narration under it.
How many seconds fit in 10 MB
The size cap becomes a length cap once you pick a format. Taking 10 MB as 10,000,000 bytes, a mono 44.1 kHz 16-bit WAV runs 88,200 bytes per second, so about 113 seconds fit. A 128 kbps MP3 runs 16,000 bytes per second, so about 625 seconds fit, which is more than the 300-second Fabric ceiling. That is simple arithmetic, not a measured Sume limit, so check your real file.
| Format | Bytes per second | Seconds in 10 MB |
|---|---|---|
| WAV, mono, 44.1 kHz, 16-bit | 88,200 | 113 |
| WAV, stereo, 44.1 kHz, 16-bit | 176,400 | 56 |
| MP3, 128 kbps | 16,000 | 625 |
The Sume path
Sume TTS defaults to MP3 at 44.1 kHz and 128 kbps, so a short line is well under the cap. Pick the voice with avatar_id or avatar_handle and Sume resolves that avatar's ready TTS voice. A Sume TTS request costs $0.0475 per 1,000 characters at list, so a 600-character line is $0.0285.
If you need the avatar to speak in a MAI voice, there is no route for it. Keep the MAI audio for other uses, for example a voice-only deliverable, and ask the avatar to speak the same text in a Sume voice.
Check before you render
Before you spend on a render, measure the clip you plan to send. The route's catalog text says the audio must be under 10 MB, so convert to MP3 at 128 kbps if a WAV is too big.
Keep the line under the 45-second Flash ceiling only if you plan to compare against MAI Flash. On Sume the avatar route, not the TTS model, sets the length limit.
There is also a rights and consent angle. A cloned or reference-matched voice from another service comes with that service's terms, and Microsoft's launch note mentions consent guardrails for cloning across languages. A presenter voice on Sume is tied to a workspace avatar whose voice status is ready, which keeps the voice, the face and the script in one place you can audit later. Whichever route you pick, write down which service made the audio before the clip ships.
For a quick decision: if the viewer must see a face speaking, make the audio on Sume. If the viewer only hears a voice, a MAI file is fine, and Sume STT or the caption routes can still read it from a public HTTPS URL.
Why this costs little
Cost sits on both sides of the choice. A Sume TTS line of 600 characters is $0.0285 at the list rate, and the talking-still render is billed on top, per second of audio. Reusing a Microsoft file would save the 2.85 cents but not the render. The saving from bringing your own audio is small, and a mismatched pipeline can cost more in rework than that.
Sources
Related posts
More in Media tools
- Match voice emotion and music mood: one mood word, two Sume fields
Set generation_config.emotion on TTS and the emotion axis in the Music prompt from the same mood word, so a short's voice and bed do not argue.
- Microsoft's Content Provenance Detection: what you can check on a file
Foundry has a detection website and API for provenance. What the page says it checks, its limits, and how to use it on an AI clip or image from any generator.
- Microsoft: C2PA may not survive a crop, transcode or compression
Microsoft's provenance page lists the edits that can drop credentials and watermarks. Which of them a Sume trim, filter or caption job could be, and a check.
- Fit TTS audio to the 5 to 14.8 second H3 Max lip-sync window
Sume's H3 Max lip-sync accepts audio of 5 to 14.8 seconds and never clamps. Plan scripts of about 63 to 197 characters, and send anything else to Fabric.
Written by Sume