MAI-Voice-2.1-Flash '60% cheaper': the baseline that implies
Microsoft calls Flash about 60% cheaper than comparable models at $15 per 1M characters. That implies a $37.50 baseline. A check against a $47.50 TTS rate.

Microsoft's launch post says MAI-Voice-2.1-Flash is about 60 percent cheaper than comparable models at $15 per 1M characters, with 55 percent faster inference (read 2026-10-03). Do the algebra and the unnamed baseline comes out at $37.50 per 1M characters, because $15 is 40 percent of $37.50. That is below Sume's $47.50 per 1M for Sonic voices, and above some rates other vendors publish.
A percentage with no named baseline is a claim you can only verify by running your own numbers. Here is how, for price and speed.
What the pages say, and what they do not
The model page lists $22 per 1M for MAI-Voice-2.1 and $15 per 1M for Flash, with about 550 ms and about 45 ms of model inference (MAI-Voice-2.1, read 2026-10-03). The launch post adds 150 ms end to end for 45 seconds of audio, 55 percent faster inference and about 60 percent cheaper than comparable models (Microsoft AI, read 2026-10-03). The pages I read do not name the comparable models. The 55 percent figure cannot be derived from the two inference numbers on the model page: 45 ms against 550 ms is a much larger gap, so it must compare against something else. Neither number is checkable from the pages alone.
The arithmetic
Flash is 31.8 percent cheaper than its own sibling, which is real and checkable. The 60 percent figure depends on a baseline you cannot see, so it is better read as 'cheap' than as a number to put in a business case.
| Quantity | Value | How |
|---|---|---|
| Flash price | $15.00 per 1M | Microsoft |
| MAI-Voice-2.1 price | $22.00 per 1M | Microsoft |
| Flash vs 2.1 | 31.8% cheaper | (22 - 15) / 22 |
| Implied 'comparable' price | $37.50 per 1M | 15 / (1 - 0.60) |
| Sume TTS Router | $47.50 per 1M | List 38 micros per character x 1.25 |
| Sume vs implied baseline | 26.7% higher | 47.5 / 37.5 - 1 |
Check any claim with one script
Paste the claimed discount and the vendor's price, and it prints the baseline the claim implies. Compare that baseline with prices you can name.
def implied_baseline(price, claimed_discount):
# price is what you pay; discount is a fraction like 0.60
return price / (1 - claimed_discount)
def pct_cheaper(cheap, dear):
return (dear - cheap) / dear * 100
flash, v21, sume = 15.0, 22.0, 47.5
base = implied_baseline(flash, 0.60)
print("implied baseline: $%.2f per 1M" % base)
print("flash vs 2.1: %.1f%% cheaper" % pct_cheaper(flash, v21))
print("flash vs sume: %.1f%% cheaper" % pct_cheaper(flash, sume))
print("sume vs implied baseline: %+.1f%%" % ((sume / base - 1) * 100))
Where else the percentage trick applies
The same algebra works on any 'X percent cheaper' claim. If a vendor says 40 percent cheaper than the leading model, divide its price by 0.60. If it says 2x faster, the baseline is twice the time. Write the implied baseline in your sheet next to the real prices you can check, and see whether any named competitor actually sits there. If none does, the claim is a marketing range, not a measurement.
It also helps to separate price per character from price per finished minute. At about 900 characters per minute, Flash is under 1.4 cents a minute and Sume's router about 4.3 cents a minute, before rounding. Both are small next to editing time, so for most teams the deciding factor is whether the voice needs a retake, how the file is delivered and how it fits the rest of the pipeline.
Decide on your own mix
Price per character is only part of cost. For ten thousand 450-character clips a month, 4.5 million characters, Flash is $67.50 and Sume is $213.75. That gap is real. But a pipeline also pays for what happens to the file: Sume's captions are $0.20 per job for videos up to 60 seconds, joins and detaches are $0.01 per job, and a render has its own price. Put the voice in the context of the whole video.
Then listen. A cheaper voice that needs a second take costs more than the dearer one that works once. Run the same ten lines on each engine and count retakes, not just dollars.
Sources
Related posts
More in Pricing
- H3 Max 1080p, 10 seconds: $1.84 before the margin change, $2.00 now
Sume bills MiniMax H3 Max at list x 1.25 like every other model, no longer x 1.15. A 10-second 1080p clip moves from $1.84 to $2.00. 480p and 768p rows too.
- MiniMax H3 Max 40% off ends Oct 15: what Sume bills
fal shows H3 Max 40% off until Oct 15. Sume bills the regular list x 1.25 throughout: $0.0625, $0.10 and $0.20 per second at 480p, 768p and 1080p.
- Balance needed for one AI generation call: published max per endpoint
Sume's catalog publishes estimated, minimum and maximum cents per endpoint. The worst-case hold for one call runs from 3 cents to $56.25 before a 402 can fire.
- Monthly AI video budget: 60 five-second clips per month on Sume
Sixty 5 second clips a month cost $30.00 on H3 Max 768p, $37.80 on Wan 3.0 or Gemini Omni Flash 720p, $113.40 on Seedance 2 and $173.40 on Seedance 2.5.
Written by Sume