Best API for text to speech: six checks to run before you pick
Compare TTS APIs on price per million characters, request cap, latency, languages, controls and word timing, with MAI-Voice-2.1 and Sume TTS 1.0 numbers.

"Best" depends on what you are voicing. A 25-second ad, a 40-minute course and a live phone agent each rank the same vendors differently. Use six checks, and score every API on the same ones. Below, Microsoft's October 1, 2026 MAI-Voice-2.1 launch numbers sit next to Sume TTS 1.0 so you can see how the checks play out.
The six checks
- Price per million characters, and how it is rounded.
- Per-request limit: characters or seconds of audio.
- Latency model: streaming first byte, or finished file.
- Language coverage, in the vendor's own list.
- Controls you actually need: speed, volume, emotion, pronunciation.
- Word timing, for captions and slicing.
The numbers
Microsoft's post (read 2026-10-05) gives MAI-Voice-2.1 at $22 per million characters in 23 languages, and the Flash variant at $15 per million, 45 seconds of audio per request and about 150 ms end-to-end latency. For Sume, TTS 1.0 is $0.0475 per 1,000 characters, or $47.50 per million, up to 20,000 characters per request, delivered as an async job result. Cartesia's page (read 2026-10-05) lists 44 languages for Sonic 3.6, which Sume TTS 1.0 runs.
| Check | MAI-Voice-2.1 | MAI-Voice-2.1-Flash | Sume TTS 1.0 |
|---|---|---|---|
| Price per 1M chars | $22 | $15 | $47.50 |
| Per-request cap | Not stated in the post | 45 s of audio | 20,000 chars, 1,200 s of audio |
| Latency model | Not stated in the post | About 150 ms end to end | Async job, finished file |
| Languages | 23 | Not stated in the post | Sonic 3.6 list, 44 |
| Word timings | Not stated in the post | Not stated in the post | timestamps.words |
How to read it
On price alone, Sume is the most expensive row. If your product is a voice that talks back inside a call, the 150 ms figure is the check that decides it, and a finished-file API is the wrong shape. If your product is finished media, such as ads, course lessons and captions, the checks that matter are the cap, the word timings and the retry behavior. The 20,000-character cap means a ten-minute script fits in one request, while a 45-second cap means about 14 stitched requests for the same script.
Do the arithmetic on your volume
Multiply your monthly characters by each price. A 300,000-character month is $6.60 at $22 per million, $4.50 at $15, and $14.25 at Sume's $47.50. Then add the cost of the stitching code, the retries and the caption pass. For a Sume job, each request also rounds up to a whole cent, which matters only for many tiny requests. The Sume router catalog post shows how to read the live price.
What this post does not claim
Microsoft did not publish a Sume-style per-request limit for the full model in the page we read, and we do not assume one. Check the vendor's current docs before you build.
Sources
Related posts
More in Comparisons
- Beyond Presence 50 credits a minute vs Sume lip sync per second
Beyond Presence bills live speech-to-video at 50 credits a minute. Sume bills a rendered lip-sync clip per second. Here is how to compare the two units.
- bitHuman local .imx or cloud avatar vs a hosted Sume avatar job
bitHuman renders avatars locally from .imx files or in its cloud, which changes your CPU bill. A hosted Sume avatar job moves the render off your agent.
- Can an AI avatar see me on a video call? Griffin perception vs Sume
Tavus says Griffin perceives visual context: 3.73 of 5 on VideoFDB perception vs 4.20 for humans. Sume avatar clips cannot see you, but a clip can be inspected.
- Can gpt-live-transcribe take an uploaded file? Endpoints vs Sume STT
OpenAI lists gpt-live-transcribe at $0.017/min and only for the realtime transcription endpoint. For a file, Sume STT is $0.01/min, up to 10 minutes per job.
Written by Sume