Audio description for 200 videos at 400 characters: TTS cost
Voicing a 400-character description track for each of 200 videos is 80,000 characters: $1.20 on MAI Flash, $1.76 on MAI-Voice-2.1, $3.80 on Sume TTS.

An audio description track of 400 characters for each of 200 videos is 80,000 characters. At list that is $1.20 on MAI-Voice-2.1-Flash, $1.76 on MAI-Voice-2.1 and $3.80 on Sume TTS, so the whole library is a few dollars of speech whichever you pick.
The bill
The money is not the hard part of audio description. Placement is: each spoken line has to land in a gap of the original dialogue and end before the next line of speech starts.
| Option | Characters | List cost |
|---|---|---|
| MAI-Voice-2.1-Flash ($15 per 1M characters) | 80,000 | $1.20 |
| MAI-Voice-2.1 ($22 per 1M characters) | 80,000 | $1.76 |
| Sume TTS ($0.0475 per 1,000 characters) | 80,000 | $3.80 |
Getting timed lines
A description track is a set of timed lines, not one long read. On Sume, split each video's description into sentences and ask for segmentation: {mode: "sentence", emit_audio: true} with a wav container, and every sentence comes back with its own audio_url and a duration. You know exactly how long each line runs before you place it.
A 400-character description is about 25 seconds of speech at a normal pace, so 40 tracks add up to about 1,000 seconds, just under the 1,200-second audio limit. Batch 40 videos per job, which is 16,000 characters, and 200 videos need five jobs. Spend is tracked per job, and every job takes its own idempotency key.
Fitting lines into the gaps
Place the lines with Timeline 1.0. A render takes one audio spine or up to 20 gapless parts, and the video clips sit on a timeline whose starts you set yourself. If a description line is longer than its gap, shorten the text, or raise generation_config.speed, which accepts values up to 1.5.
Sume's duration check protects you from a long line. A job that comes out over 1,200 seconds fails with tts_duration_exceeded, which a 400-character line never approaches. The practical limit is the gap in the picture, not the platform.
Flash makes up to 45 seconds of audio per call, which also covers one description line. The comparison again comes down to what happens after the call: with Sume the line is already a hosted file that a render can take.
What the table leaves out
Microsoft's figures are the launch prices from 2026-10-01 (Microsoft AI, read 2026-10-05): $22 per 1M characters for MAI-Voice-2.1 and $15 per 1M characters for MAI-Voice-2.1-Flash, which makes up to 45 seconds of audio per call. Sume's figure is its list rate per transcript character, and on Sume spaces and punctuation count toward usage, so count characters of the final text.
The cheaper per-character price does not make the Microsoft route a drop-in swap. A Sume TTS request is an async job whose result carries a hosted audio_url, duration, optional word timings and sentence slices, and that hosted file is what Timeline audio and renders accept. Count the extra work of storing and joining files outside Sume before you call the gap a saving.
Check language coverage first. Microsoft lists 23 languages and 26 locales for MAI-Voice-2.1, and Sume takes a language field that you must set for every non-English transcript. A cheaper model that does not speak your target language is not an option, so settle that before you compare totals.
Pilot and keep one voice
Pilot on five videos. Read each result's segments[] durations, compare them with the gaps you measured in the video, and mark any line that overruns. Fix the text, not the voice speed, wherever you can, because slower or faster speech is the first thing a listener who relies on description will notice.
Keep one voice for the whole library so the track sounds the same on every title, and store the voice id and speed with the project. Word timings are an option in the request, so ask for them when a transcript must ship next to the audio.
Sources
Related posts
More in Pricing
- Why your Sume balance dropped for a job that is still queued
Sume reserves the estimated cost when it accepts a job, not when it starts. Six queued 10-second Wan 720p jobs hold $7.50 at once. How to read it.
- Black Friday hook tests: 20 ten-second variants on seven Sume rates
What 20 ten-second Black Friday ad variants cost on Wan 3.0, Gemini Omni Flash 1.1 and MiniMax H3 Max at Sume's per-second rates, from $12.50 to $50.00.
- Budget for 12 UGC ad variants: draft on Mini, finish on Seedance 2.5
What 12 vertical UGC ad variants cost on Sume when you draft on seedance-2-mini at 480p and finish the winners on seedance-2.5 at 720p. Full arithmetic inside.
- Caption a 10-minute TikTok: ten 60-second jobs, $3.30 with Sume
Sume states the caption price for videos up to 60 seconds. Cut a 10-minute TikTok into ten 60-second parts: $3.30 with inspect, trim, captions and reassembly.
Written by Sume