MAI-Voice-2.1-Flash makes 45 s of audio: narrating a 5-minute script
Microsoft lists 45 seconds of audio per Flash generation. A 5-minute script needs about seven chunks. A Python splitter, and how Sume handles long text.

A five-minute narration does not fit in one MAI-Voice-2.1-Flash call. Microsoft's announcement says Flash can generate 45 seconds of audio with about 150 ms end-to-end latency (read 2026-10-03), so a 300-second script is at least seven separate generations, and you have to cut the text at sentence boundaries yourself and join the audio afterwards.
This post shows the arithmetic, a splitter you can run today, and what changes if you send the same script to Sume's TTS Router, which takes up to 20,000 characters in a single job.
What Microsoft actually says about length
The 45-second figure appears in the launch post for the new audio models, next to the 150 ms claim and the $15 per 1M characters price for Flash (Microsoft AI, read 2026-10-03). The MAI-Voice-2.1 model page lists the two models side by side at $22 and $15 per 1M characters, with about 550 ms and about 45 ms model inference (model page, read 2026-10-03). The pages I read do not state a character limit per request, so treat the 45 seconds as the working ceiling and size your chunks from it.
Speech rate decides how many characters fit in 45 seconds. This post uses a planning assumption, not a Microsoft number: about 15 characters of spoken text per second, which is roughly 150 words a minute at six characters a word including the space. At that rate 45 seconds is about 675 characters. Cutting at 600 leaves headroom for slower voices and pauses.
Chunk budget for common scripts
The cost does not change with chunking, because the price is per character. What changes is the number of requests, retries and seams. Every seam is a place where intonation can reset, so cut on full sentences and keep each chunk's last sentence short.
| Script length | Characters (approx.) | Flash chunks at 600 | Cost at $15 per 1M |
|---|---|---|---|
| 30-second ad | 450 | 1 | $0.0068 |
| 2-minute explainer | 1,800 | 3 | $0.027 |
| 5-minute walkthrough | 4,500 | 8 | $0.068 |
| 20-minute lesson | 18,000 | 30 | $0.27 |
A splitter that cuts on sentence ends
This runs as written. It never splits inside a sentence unless one sentence alone is longer than the limit, in which case it falls back to the clause comma.
import re
def chunk_script(text, limit=600):
sentences = re.split(r"(?<=[.!?])\s+", text.strip())
chunks, current = [], ""
for sentence in sentences:
if len(sentence) > limit:
# fall back to commas for a runaway sentence
parts = re.split(r"(?<=,)\s+", sentence)
else:
parts = [sentence]
for part in parts:
if current and len(current) + 1 + len(part) > limit:
chunks.append(current)
current = part
else:
current = f"{current} {part}".strip()
if current:
chunks.append(current)
return chunks
script = " ".join(["This is sentence number %d of the walkthrough." % i for i in range(1, 121)])
pieces = chunk_script(script)
print(len(script), "characters ->", len(pieces), "chunks")
print("longest chunk:", max(len(p) for p in pieces))
print("Flash cost: $%.4f" % (len(script) * 15 / 1_000_000))
The same script on Sume
Sume's TTS Router accepts a transcript of 1 to 20,000 characters per job, and the model field picks a Cartesia Sonic id from the catalog (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). A 4,500-character walkthrough is therefore one job, not eight. The price is the provider list rate times 1.25, which works out to $47.50 per 1M characters, so the 4,500-character script costs about $0.21 against $0.068 on Flash. You pay more per character and skip the splitting and the seams.
Sume's TTS Router is not streaming and is not a 150 ms product: you submit a job, poll /v1/jobs/:id/status, and read the audio from /result (Jobs and results). That fits narration you render ahead of time. It does not fit a voice agent that must start speaking while the user is still finishing a sentence.
If you do need to join separately generated takes into one file, timeline audio concatenates up to 20 Sume-hosted audio parts sample by sample for $0.01 flat per job, with no silence at the seams. That covers audio made on Sume. Audio from another vendor has to be imported onto the Sume media host first, per the same page.
Decision
Choose Flash when speed and the per-character price matter and your lines are short: prompts, replies, captions read aloud. Choose a single long job when the script is one continuous read and you would rather not manage seams. Whichever you pick, count characters before you submit and keep a cap on the job cost.
One practical habit helps with either route: generate the first chunk and the last chunk of a long script first and listen to them side by side. If the voice, pace and loudness of the two ends match, the middle will almost always match too, and you have not spent the whole budget finding out. If they drift, shorten the chunks, keep the same voice selector and language on every call, and regenerate only the chunk that sounds off rather than the whole script. Per-character pricing makes that cheap: replacing one 600-character chunk of a 4,500-character script costs under a cent at Flash's rate and about three cents at Sume's.
- Under 45 seconds per line: one Flash call, no splitting.
- Over 45 seconds: split at sentence ends near 600 characters, generate, join.
- Over 600 characters but one voice and one take wanted: a single TTS Router job up to 20,000 characters.
Sources
Related posts
More in Models
- MAI-Voice-2.1 on OpenRouter and Vercel: is it in Sume's TTS Router?
Microsoft lists MAI voices on Foundry, OpenRouter and Vercel. Sume's TTS Router lists only Cartesia Sonic ids today. How to check the live catalog.
- MiniMax H3 2K and 4K upscale on Sume: H3 accepts them, H3 Max does not
minimax-h3 will price a 2K or 4K request even though its resolution list shows only 480p and 768p. minimax-h3-max rejects both. Costs for 5 to 15 seconds.
- MiniMax H3 or H3 Max on Sume: which id for a 5 to 15 second clip?
Choose minimax-h3 for a 480p or 768p draft at $0.075 a second, minimax-h3-max when you need 1080p. Same 5-15 s range, same references; Max costs more at 768p.
- MiniMax H3 reference-to-video: five images free, then $0.08 each
On the minimax-h3 id, reference-to-video adds a list fee for every reference image after the fifth. minimax-h3-max does not. Worked 10-second totals.
Written by Sume