MAI-Voice-2.1 audio in a Sume Timeline: 9 clips, 261 s
Sume does not list MAI-Voice. To use its audio in a Timeline, import the files, concat at $0.01, render 261 s as 5 minutes ($0.50). Read 2026-10-08.

Sume does not list any MAI model, but you can still use MAI-Voice-2.1 audio in a Sume video: generate the lines at Microsoft, import the files so they live on media.sume.com, join them with one Timeline audio concat job at $0.01, and render the video. For nine clips totaling 261 seconds, the render bills 5 started minutes, $0.50.
The cost stack for a 261-second example
The example is 9 lines totaling 3,950 characters. Microsoft's page prices are $22 per million characters for MAI-Voice-2.1 and $15 for Flash. Timeline audio concat is $0.01 per job and Timeline render is $0.10 per started minute, so 261 seconds is ceil(261 / 60) = 5 minutes. The cost of importing files is not stated in the docs pages I read, so it is not in the table.
| Step | Basis | Cost |
|---|---|---|
| Speech, MAI-Voice-2.1 | 3,950 chars x $22 / 1M | $0.0869 |
| Speech, MAI-Voice-2.1-Flash | 3,950 chars x $15 / 1M | $0.05925 |
| Concat 9 files | 1 job x $0.01 | $0.01 |
| Render 261 s | 5 started minutes x $0.10 | $0.50 |
| Total with Flash | $0.56925 | |
| Total with 2.1 | $0.5969 |
Rules the Timeline docs impose
Every audio URL passed to Timeline audio must be this workspace's media.sume.com audio, so off-host links are refused at admit and you must import first (the docs name POST /v1/media-imports). All parts of a concat must share a channel layout, otherwise the job fails with audio_parts_channel_mismatch, so export all nine MAI files as the same mono or stereo layout. Concat takes 1 to 20 parts, so nine fit in one job. Keep wav if the audio will be joined again; mp3 adds priming padding at every edge.
The result lists segments[] with start offsets for each part. Use those to place the nine video clips so each picture starts on its line.
What this does not give you
Sume's TTS job records the voice, model and synthesis settings so a later line can match; audio you bring from another provider has no such record in Sume, so keep the provider's voice id and settings next to the project yourself. The price comparison also changes at small sizes: the same 3,950 characters as nine Sume TTS jobs of about 439 characters each is $0.1876 at list price, and the Sume path needs no import step.
The order of calls
Import each of the nine files (the Timeline audio docs name POST /v1/media-imports), collect the media.sume.com URLs, then call POST /v1/timeline-1.0/audio with operation concat and a parts[] array of nine objects. Send an Idempotency-Key, poll GET /v1/jobs/:id/status, read audio_url and segments[] from the result, and pass audio_url as audio.url on the render with audio.duration_seconds of 261. The render reserve is ceil(261 / 60) = 5 minutes, which is $0.50, so trimming 21 seconds of tail to reach 240 seconds would cut the render to 4 minutes ($0.40).
Costs on the Sume side are therefore $0.01 for the concat and $0.50 for the render, $0.51 in total, and nothing in the docs I read prices the import step, so this post does not claim a figure for it. The MAI-Voice-2.1 speech itself is billed by Microsoft at $22 per million characters.
Sources
Related posts
More in Developers
- Validate a Gemini Omni Flash 1.1 request in Python before sending
A 28-line Python check for Sume's gemini-omni-flash-1.1 rules: 3-10 s, 10 reference images, 3 reference videos, no audio off, and edit mode exclusions.
- A Veo 3.1 call becomes a Sume Omni job in under 30 lines of Python
Replace a Veo 3.1 request with a Sume gemini-omni-flash-1.1 job: submit, poll every 30 seconds, download the mp4 and read usage.cost. Standard library only.
- Veo 3.1 previews end in 14 days: a dated checklist, Oct 8 to Oct 22
Google's three Veo 3.1 preview ids shut down on October 22, 2026. A day-by-day checklist from today, with the Sume model id and limits to test against.
- Vercel 800 s max duration: do you still need a Sume webhook?
Vercel Pro allows 800 s functions and a 30-minute beta. A Sume video job can still outlast one request, so use async or webhook mode and return in seconds.
Written by Sume