MAI-Voice-2.1-Flash tops out at 45 seconds: what for longer?

MAI-Voice-2.1-Flash is reported at up to 45 s of audio and 150 ms latency. Longer lines need the standard model or a TTS taking 20,000 characters.

5 min readSume
All posts

Unite.ai reports MAI-Voice-2.1-Flash at up to 45 seconds of audio with 150 ms end-to-end latency, at $15 per 1M characters. Anything longer than 45 seconds should use a different model. Microsoft's MAI page says the long-form model prioritises naturalness over latency, so it is not for live use.

Which one for which job

MAI voice models as reported (read 2026-10-04)
ModelPrice per 1M charactersReported behaviour
MAI-Voice-2.1$22Long-form, naturalness over latency
MAI-Voice-2.1-Flash$15Up to 45 s of audio, 150 ms end to end

Sume's limit is in characters

Sume's TTS accepts up to 20,000 characters per request and prices at $0.0475 per 1,000. Spoken length depends on pace, so a 20,000-character script is a long read; the speed field takes slow, normal or fast.

  • Short prompts under 45 seconds are within Flash range.
  • Longer reads: split, generate, then join with Timeline audio concat.
  • Concat output is limited to 1,800 seconds.

Worked example

A 90-second voice line at about 15 characters per second is roughly 1,350 characters: $0.0641 on Sume at $0.0475 per 1,000, against $0.0203 on Flash at $15 per million. The speaking rate is an assumption.

Preview status

The MAI voices page marks the service as public preview with no SLA and not recommended for production.

Sources

Related posts

More in Models

All Models posts

Written by Sume