MAI-Voice-2.1 and Format runs: get the voiceover as its own file

A Format run cannot be told to use MAI-Voice-2.1, since its tools pick the audio model. Bind an audio field to get the voiceover track as a file.

4 min readSume
All posts

No, you cannot point a Sume Format run at Microsoft's MAI-Voice-2.1. The model field on a run selects only the orchestrating LLM, and the Format's tools select the image, video and audio models. What you can do is ask for the voiceover as a separate file, by binding an audio field to SumeMediaFile#, and then check its length against the video.

What Microsoft launched

Microsoft's model page lists two models. MAI-Voice-2.1 is quoted at about 550 ms of model latency and $22 per 1M characters, and MAI-Voice-2.1-Flash at about 45 ms and $15 per 1M characters. It covers 23 languages and is pitched at voice-over and audiobooks for the standard model and at call-center agents for Flash. Those are Microsoft's numbers for Microsoft's API. They say nothing about what a Sume Format run calls internally.

Why a Format user cares

A rendered ad does not need a 45 ms model, because nobody is waiting on a live conversation. What an ad pipeline needs is the voice track as an asset it can re-time, caption or swap. A Format run mixes the voiceover into its video, so unless you ask, you only see the final cut. Asking is a schema decision.

Ask for the audio as a field

SumeMediaFile has type, url, content_type, duration_ms and more. When the run reports a duration_ms, Sume compares it with the ledger of what the run made, and a value more than 10% off fails the projection. So the number you read back is the artifact's length or it is unverified, never a guess. A nullable union keeps a Format that makes no standalone audio from failing.

Fields for a voiceover-aware output schema (Sume docs read 2026-10-05, Microsoft page read 2026-10-05)
FieldTypeReads as
videoSumeMediaFileThe assembled ad; also name it in primary_output_key
voiceoverSumeMediaFile or nullStandalone audio, if the recipe generates it
voiceover_languagestring or nullWritten by the run; do not expect your input value back
script_usedstring or nullThe final text the run spoke
{
  "output_schema": {
    "name": "acme/ad-with-voice/v1", "strict": true,
    "schema": {
      "type": "object", "additionalProperties": false,
      "required": ["video", "voiceover", "script_used"],
      "properties": {
        "video": {"$ref": "SumeMediaFile#"},
        "voiceover": {"anyOf": [{"$ref": "SumeMediaFile#"}, {"type": "null"}]},
        "script_used": {"type": ["string", "null"]}
      }
    }
  },
  "primary_output_key": "video"
}

Where to put the script

Put the script and the language in input, not in the instruction. Sume writes input to a file in the run's workspace and tells the agent it is data. On the projection fallback path the schema pass sees only the run's media and closing text, so a language code you sent may come back null. Keep it in your own record, keyed by the run id or your idempotency key.

When the field is null

If the receipt shows voiceover: null, the recipe did not make a separate track. The audio is inside the video, and artifacts[] is the place to look for anything else the run generated. If you need MAI-Voice-2.1 specifically, call Microsoft directly and lay the file over the video with a timeline step. Do not expect a Format run to route there.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume