MAI-Voice-2.1 Flash narration for a video cut

Narration from any voice model can be the audio spine of a Sume Timeline cut. Import the file first, then render at $0.10 per output minute.

4 min readSume
All posts

To put narration under a video cut on Sume, make the audio file anywhere, import it into your workspace so it has a media.sume.com URL, and pass it as the audio.url spine of a Timeline 1.0 render. Timeline then lays ordered video slots over that audio and returns one MP4.

What was reported on Oct 1, 2026

The Digital Applied October 2026 tracker lists Microsoft's MAI-Voice-2.1 and a Flash variant. Figures below are as that page states them; this post has not measured them.

MAI-Voice-2.1 figures as reported (read 2026-10-03)
ItemMAI-Voice-2.1MAI-Voice-2.1-Flash
Price per million characters$22$15
Speed claimBaselineVendor claim: 55% faster, roughly 60% cheaper than comparable models
Latency claimNot stated45 s of audio in 150 ms end to end
Languages2323

Where narration goes in a Sume cut

Timeline 1.0 takes one audio spine plus ordered video[] slots. Required fields are audio.duration_seconds (1 to 1800) and either audio.url or audio.parts[], plus 1 to 200 video slots. Every URL must already be a media.sume.com artifact or asset in your workspace; off-host URLs are rejected at admit, so import first with POST /v1/media-imports.

The docs do not list MAI-Voice as a Sume model, so treat it as an outside source of audio. For narration made on Sume, text-to-speech jobs record the voice and model they used.

Several narration takes

If narration comes as several files, there are two options. Put them in audio.parts[] (up to 20 gapless slices) when the join is needed only inside one render. Use Timeline audio with operation: "concat" when you want a reusable merged file. The join is sample-domain, with no re-synthesis and no silence at the seams.

The concat result returns segments[] with start offsets, which you use to align each video[].start.

Cost and checks

The public Timeline rate is $0.10 per ceil(output minute), with the reserve equal to ceil(audio.duration_seconds / 60) minutes. Confirm live pricing in GET /v1/catalog. POST /v1/timeline-1.0/plan compiles the document and estimates cost without creating a job.

Narration speed or latency from the voice vendor does not change what Timeline costs; it only changes how soon the audio file exists.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume