MAI-Voice-2.1 audio in a Sume Timeline: 9 clips, 261 s

Sume does not list MAI-Voice. To use its audio in a Timeline, import the files, concat at $0.01, render 261 s as 5 minutes ($0.50). Read 2026-10-08.

5 min readSume
All posts

Sume does not list any MAI model, but you can still use MAI-Voice-2.1 audio in a Sume video: generate the lines at Microsoft, import the files so they live on media.sume.com, join them with one Timeline audio concat job at $0.01, and render the video. For nine clips totaling 261 seconds, the render bills 5 started minutes, $0.50.

The cost stack for a 261-second example

The example is 9 lines totaling 3,950 characters. Microsoft's page prices are $22 per million characters for MAI-Voice-2.1 and $15 for Flash. Timeline audio concat is $0.01 per job and Timeline render is $0.10 per started minute, so 261 seconds is ceil(261 / 60) = 5 minutes. The cost of importing files is not stated in the docs pages I read, so it is not in the table.

Cost of the example, read 2026-10-08 (import step not included)
StepBasisCost
Speech, MAI-Voice-2.13,950 chars x $22 / 1M$0.0869
Speech, MAI-Voice-2.1-Flash3,950 chars x $15 / 1M$0.05925
Concat 9 files1 job x $0.01$0.01
Render 261 s5 started minutes x $0.10$0.50
Total with Flash$0.56925
Total with 2.1$0.5969

Rules the Timeline docs impose

Every audio URL passed to Timeline audio must be this workspace's media.sume.com audio, so off-host links are refused at admit and you must import first (the docs name POST /v1/media-imports). All parts of a concat must share a channel layout, otherwise the job fails with audio_parts_channel_mismatch, so export all nine MAI files as the same mono or stereo layout. Concat takes 1 to 20 parts, so nine fit in one job. Keep wav if the audio will be joined again; mp3 adds priming padding at every edge.

The result lists segments[] with start offsets for each part. Use those to place the nine video clips so each picture starts on its line.

What this does not give you

Sume's TTS job records the voice, model and synthesis settings so a later line can match; audio you bring from another provider has no such record in Sume, so keep the provider's voice id and settings next to the project yourself. The price comparison also changes at small sizes: the same 3,950 characters as nine Sume TTS jobs of about 439 characters each is $0.1876 at list price, and the Sume path needs no import step.

The order of calls

Import each of the nine files (the Timeline audio docs name POST /v1/media-imports), collect the media.sume.com URLs, then call POST /v1/timeline-1.0/audio with operation concat and a parts[] array of nine objects. Send an Idempotency-Key, poll GET /v1/jobs/:id/status, read audio_url and segments[] from the result, and pass audio_url as audio.url on the render with audio.duration_seconds of 261. The render reserve is ceil(261 / 60) = 5 minutes, which is $0.50, so trimming 21 seconds of tail to reach 240 seconds would cut the render to 4 minutes ($0.40).

Costs on the Sume side are therefore $0.01 for the concat and $0.50 for the render, $0.51 in total, and nothing in the docs I read prices the import step, so this post does not claim a figure for it. The MAI-Voice-2.1 speech itself is billed by Microsoft at $22 per million characters.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume