MAI-Voice-2.1-Flash makes 45 s of audio: narrating a 5-minute script

Microsoft lists 45 seconds of audio per Flash generation. A 5-minute script needs about seven chunks. A Python splitter, and how Sume handles long text.

5 min readSume
All posts

A five-minute narration does not fit in one MAI-Voice-2.1-Flash call. Microsoft's announcement says Flash can generate 45 seconds of audio with about 150 ms end-to-end latency (read 2026-10-03), so a 300-second script is at least seven separate generations, and you have to cut the text at sentence boundaries yourself and join the audio afterwards.

This post shows the arithmetic, a splitter you can run today, and what changes if you send the same script to Sume's TTS Router, which takes up to 20,000 characters in a single job.

What Microsoft actually says about length

The 45-second figure appears in the launch post for the new audio models, next to the 150 ms claim and the $15 per 1M characters price for Flash (Microsoft AI, read 2026-10-03). The MAI-Voice-2.1 model page lists the two models side by side at $22 and $15 per 1M characters, with about 550 ms and about 45 ms model inference (model page, read 2026-10-03). The pages I read do not state a character limit per request, so treat the 45 seconds as the working ceiling and size your chunks from it.

Speech rate decides how many characters fit in 45 seconds. This post uses a planning assumption, not a Microsoft number: about 15 characters of spoken text per second, which is roughly 150 words a minute at six characters a word including the space. At that rate 45 seconds is about 675 characters. Cutting at 600 leaves headroom for slower voices and pauses.

Chunk budget for common scripts

The cost does not change with chunking, because the price is per character. What changes is the number of requests, retries and seams. Every seam is a place where intonation can reset, so cut on full sentences and keep each chunk's last sentence short.

Planning math from the assumption above (15 characters per second, 600-character chunks), vendor limit read 2026-10-03.
Script lengthCharacters (approx.)Flash chunks at 600Cost at $15 per 1M
30-second ad4501$0.0068
2-minute explainer1,8003$0.027
5-minute walkthrough4,5008$0.068
20-minute lesson18,00030$0.27

A splitter that cuts on sentence ends

This runs as written. It never splits inside a sentence unless one sentence alone is longer than the limit, in which case it falls back to the clause comma.

import re

def chunk_script(text, limit=600):
    sentences = re.split(r"(?<=[.!?])\s+", text.strip())
    chunks, current = [], ""
    for sentence in sentences:
        if len(sentence) > limit:
            # fall back to commas for a runaway sentence
            parts = re.split(r"(?<=,)\s+", sentence)
        else:
            parts = [sentence]
        for part in parts:
            if current and len(current) + 1 + len(part) > limit:
                chunks.append(current)
                current = part
            else:
                current = f"{current} {part}".strip()
    if current:
        chunks.append(current)
    return chunks

script = " ".join(["This is sentence number %d of the walkthrough." % i for i in range(1, 121)])
pieces = chunk_script(script)
print(len(script), "characters ->", len(pieces), "chunks")
print("longest chunk:", max(len(p) for p in pieces))
print("Flash cost: $%.4f" % (len(script) * 15 / 1_000_000))

The same script on Sume

Sume's TTS Router accepts a transcript of 1 to 20,000 characters per job, and the model field picks a Cartesia Sonic id from the catalog (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). A 4,500-character walkthrough is therefore one job, not eight. The price is the provider list rate times 1.25, which works out to $47.50 per 1M characters, so the 4,500-character script costs about $0.21 against $0.068 on Flash. You pay more per character and skip the splitting and the seams.

Sume's TTS Router is not streaming and is not a 150 ms product: you submit a job, poll /v1/jobs/:id/status, and read the audio from /result (Jobs and results). That fits narration you render ahead of time. It does not fit a voice agent that must start speaking while the user is still finishing a sentence.

If you do need to join separately generated takes into one file, timeline audio concatenates up to 20 Sume-hosted audio parts sample by sample for $0.01 flat per job, with no silence at the seams. That covers audio made on Sume. Audio from another vendor has to be imported onto the Sume media host first, per the same page.

Decision

Choose Flash when speed and the per-character price matter and your lines are short: prompts, replies, captions read aloud. Choose a single long job when the script is one continuous read and you would rather not manage seams. Whichever you pick, count characters before you submit and keep a cap on the job cost.

One practical habit helps with either route: generate the first chunk and the last chunk of a long script first and listen to them side by side. If the voice, pace and loudness of the two ends match, the middle will almost always match too, and you have not spent the whole budget finding out. If they drift, shorten the chunks, keep the same voice selector and language on every call, and regenerate only the chunk that sounds off rather than the whole script. Per-character pricing makes that cheap: replacing one 600-character chunk of a 4,500-character script costs under a cent at Flash's rate and about three cents at Sume's.

  • Under 45 seconds per line: one Flash call, no splitting.
  • Over 45 seconds: split at sentence ends near 600 characters, generate, join.
  • Over 600 characters but one voice and one take wanted: a single TTS Router job up to 20,000 characters.

Sources

Related posts

More in Models

All Models posts

Written by Sume