MAI-Voice-2.1 Flash narration for a video cut
Narration from any voice model can be the audio spine of a Sume Timeline cut. Import the file first, then render at $0.10 per output minute.

To put narration under a video cut on Sume, make the audio file anywhere, import it into your workspace so it has a media.sume.com URL, and pass it as the audio.url spine of a Timeline 1.0 render. Timeline then lays ordered video slots over that audio and returns one MP4.
What was reported on Oct 1, 2026
The Digital Applied October 2026 tracker lists Microsoft's MAI-Voice-2.1 and a Flash variant. Figures below are as that page states them; this post has not measured them.
| Item | MAI-Voice-2.1 | MAI-Voice-2.1-Flash |
|---|---|---|
| Price per million characters | $22 | $15 |
| Speed claim | Baseline | Vendor claim: 55% faster, roughly 60% cheaper than comparable models |
| Latency claim | Not stated | 45 s of audio in 150 ms end to end |
| Languages | 23 | 23 |
Where narration goes in a Sume cut
Timeline 1.0 takes one audio spine plus ordered video[] slots. Required fields are audio.duration_seconds (1 to 1800) and either audio.url or audio.parts[], plus 1 to 200 video slots. Every URL must already be a media.sume.com artifact or asset in your workspace; off-host URLs are rejected at admit, so import first with POST /v1/media-imports.
The docs do not list MAI-Voice as a Sume model, so treat it as an outside source of audio. For narration made on Sume, text-to-speech jobs record the voice and model they used.
Several narration takes
If narration comes as several files, there are two options. Put them in audio.parts[] (up to 20 gapless slices) when the join is needed only inside one render. Use Timeline audio with operation: "concat" when you want a reusable merged file. The join is sample-domain, with no re-synthesis and no silence at the seams.
The concat result returns segments[] with start offsets, which you use to align each video[].start.
Cost and checks
The public Timeline rate is $0.10 per ceil(output minute), with the reserve equal to ceil(audio.duration_seconds / 60) minutes. Confirm live pricing in GET /v1/catalog. POST /v1/timeline-1.0/plan compiles the document and estimates cost without creating a job.
Narration speed or latency from the voice vendor does not change what Timeline costs; it only changes how soon the audio file exists.
Sources
Related posts
More in Media tools
- Mercado Libre Clips: 61 s needs two clips in Timeline
Mercado Libre Clips run 10 to 61 seconds. Some Sume models make 30 s clips, so two in Timeline give a 60 s vertical Clip for one billed minute.
- Meta Reels ads convert non-9:16 media: send your own 9:16 clip
Meta says Reels placement can apply automatic templates that convert non-9:16 media. Render the 9:16 yourself with Sume timeline fit modes and keep the framing.
- Product still 30%, clip 70%: compose stack ratio in pixels
Sume compose stack ratio is the still's share of the frame, 0.1 to 0.9, and the video takes the rest. Pixel table for 1080x1920 plus a request that sets 0.3.
- Pull a voice track from AI video, then join takes gaplessly
Extract a wav or mp3 from a Sume-hosted video with POST /v1/audio-detach, then join up to 20 audio parts with Timeline audio. Fields, caps, $0.01 rate.
Written by Sume