Sume TTS is a job, not a stream: cost and fit for a voice agent reply
A 45-second spoken reply is about 675 characters: $0.0320625 on Sume TTS, $0.010125 on MAI-Voice-2.1-Flash. Sume returns a finished file, not a live stream.

Sume TTS is not a streaming service: each request creates an asynchronous job and returns a finished audio file, and the synchronous option waits at most 30 seconds. A 45-second spoken reply of about 675 characters costs $0.0320625 on Sume (0.675 x $0.0475) and $0.010125 on MAI-Voice-2.1-Flash (675 x $15 / 1M). For a live conversation where the caller waits, use a streaming product; for narration, ads and clips where a file is the goal, the job model fits.
What the vendors say about speed
Microsoft reports that MAI-Voice-2.1-Flash returns about 45 seconds of audio in about 150 ms end to end (Microsoft AI, read 2026-10-09). I treat that as the vendor's own claim, not a measurement of mine. Sume's docs give no latency figure for TTS, only the request shape and the 30 second ceiling on waiting, so none is claimed here.
| Question | Sume TTS 1.0 | MAI-Voice-2.1-Flash |
|---|---|---|
| Delivery | Async job, finished file | Vendor describes fast end-to-end synthesis |
| Wait option | sync or subscribe, up to 30 s | Not described in text read |
| 45 s reply, 675 characters | $0.0320625 | $0.010125 |
| Per 1,000 characters | $0.0475 | $0.015 |
Where a job model works
If your product answers a person on the phone, every extra second of waiting is audible, so a job that completes and then hands over a URL is the wrong tool whatever its price. If your product writes the script first, such as a product video, a narrated lesson or an ad, the file is what you need and a job gives you retries, idempotency keys and a webhook when the file is ready.
A hybrid is common: a live layer for conversation, and Sume for the artifacts around it, such as turning a call summary into a narrated clip, or a script into a voice-over with captions.
Cost of a day of replies
One thousand 675-character replies a day is 675,000 characters. On Sume that is 675 x $0.0475 = $32.0625. On the Flash model it is 0.675 x $15 = $10.125. The gap is $21.94 a day, but only the second option is built for a conversation, as far as the pages I read say.
Using the 30 second wait
The sync option lets one request wait for the job to finish, for at most 30 seconds, and subscribe does the same over a stream of status events. If the job is not done by then, you receive the job id and poll or use a webhook. For a long narration this is fine. For a reply that must start playing in a fraction of a second it is not a streaming API, and no setting makes it one. Plan the product around that fact instead of discovering it in a demo.
Sources
Related posts
More in Comparisons
- Switching Fabric to H3 Max Lip Sync: 6 s, 12 s and 20 s clips compared
Fabric and H3 Max Lip Sync share a body on Sume, so switching is a URL change. At 6 s H3 768p is $0.60 vs Fabric 720p $1.125; at 20 s H3 Max is not eligible.
- Synthesia Basic free 10 minutes vs 10 minutes of Sume avatar video
Synthesia Basic is $0 with 10 minutes and 9 avatars; the same 10 minutes on Sume standard is $110.40, so the free plan wins until you outgrow it.
- Tavus Growth's 10 concurrent streams vs Sume's 4, 8 and 20 job slots
Tavus Growth allows 10 concurrent streams. Sume Pro runs 4 generation jobs at once, Startup 8, Scale 20, and queues the rest. What each limit actually counts.
- Tavus overage $0.37 a minute vs Sume avatar $0.184 a second: 30x gap
Tavus Starter overage is $0.37 per minute ($0.00617 a second); Sume avatar standard is $0.184 a second ($11.04 a minute). Why the products are not comparable.
Written by Sume