NVIDIA Magpie TTS for voice agents vs Sume TTS 1.0: which to use?
Magpie TTS is an open-weights, low-latency model for live voice agents in 12 languages. Sume TTS is async and non-streaming. How to choose, read 2026-10-10.

For a live voice agent that must start speaking in tens of milliseconds, Sume TTS 1.0 is the wrong tool and NVIDIA Magpie TTS is a better fit; for narration, video voiceover and any audio you render ahead of time, Sume TTS is the simpler one. Sume TTS 1.0 is async-only and non-streaming, so it returns a finished file rather than audio you can play while it is still being made.
The Magpie facts below come from NVIDIA's Hugging Face blog post of August 10, 2026, read on 2026-10-10. Sume facts come from the repository and docs.
What NVIDIA says about Magpie TTS
The post describes a 364-million-parameter open-weights text-to-speech model under the NVIDIA Open Model License, covering 12 languages: English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, Modern Standard Arabic, Korean and Brazilian Portuguese. It offers a male and a female voice per language, and code-switching for Hindi and Japanese. The post does not mention voice cloning or SSML.
The headline numbers are time to first audio: 32 ms on a B200, 47 ms on an H100, 53 ms on a DGX Spark and 79 ms on an A100. With 64 concurrent streams on a B200 the figure rises to 239 ms. These are NVIDIA's measurements on NVIDIA hardware, not something Sume has reproduced. You deploy it through NIM or the open checkpoint, on GPUs you run.
| Hardware | Time to first audio | Load |
|---|---|---|
| B200 | 32 ms | Single stream |
| H100 | 47 ms | Single stream |
| DGX Spark | 53 ms | Single stream |
| A100 | 79 ms | Single stream |
| B200 | 239 ms | 64 concurrent streams |
What Sume TTS 1.0 does
Sume TTS 1.0 takes a transcript of up to 20,000 characters and a voice, and returns an audio file through a job. Modes are async, sync, subscribe and webhook; sync and subscribe wait at most 30 seconds before you poll. There is no audio stream, so there is no time-to-first-audio number to compare, and this post does not invent one.
Pricing is $0.0475 per 1,000 characters. Synthesized audio over 1,200 seconds fails with tts_duration_exceeded, so long scripts are split. Output formats are mp3, wav or raw, with sample rates from 8,000 to 48,000 Hz, and word timestamps are available when you need captions or edit points. Voices come from avatar_id, avatar_handle or a voice id; voice cloning is app-only and not in the API.
Pick by job, not by model
The deciding fact is whether a person is waiting on the other end of a live conversation. If yes, you need a streaming engine whose first audio arrives almost immediately, and open weights on your own GPUs is a legitimate way to get it. If the audio is rendered, reviewed and then published, async is fine and you avoid running anything.
- Live phone or in-app agent: choose a streaming model such as Magpie, and budget GPUs, scaling and monitoring yourself.
- Video voiceover, explainer or ad read: choose Sume TTS and join takes with Timeline audio when you need one file.
- Compliance needs on-premises audio: open weights keep text on your network; Sume is a hosted API.
- Language outside Magpie's 12 or a specific accent: check each engine's voice list directly before committing.
A hybrid that keeps both honest
Many teams need both: a live agent for conversation and prerecorded audio for the content around it, such as a welcome message, a hold announcement or a product video. Nothing stops you from using a streaming model for the first and Sume for the second, as long as the voices are close enough that callers do not notice the handoff.
To see how Sume jobs are polled and when a result is ready, read jobs and results. For the latency question in the abstract, 150 ms claims vs async TTS goes through which jobs need which. For another open-weights comparison see Voxtral TTS.
Limits of this comparison
It does not compare naturalness, because neither source supports a ranking. It also does not price the self-hosted route: GPU hours, ops time and licence review sit with you. Re-read NVIDIA's license terms before shipping, since the Open Model License has conditions that this post does not summarize.
Sources
Related posts
More in Comparisons
- A 10-second Grok Imagine 1.5 Lite clip vs Sume 720p
xAI lists Grok Imagine Video 1.5 Lite at $0.020 a second: $0.20 for one 10-second clip. Sume's grok-imagine-video-1.5 bills $0.125 at 720p.
- A 10-second Vidu Q4 Preview clip vs Sume at 720p
Vidu Q4 Preview lists launch pricing from $0.014 a second. For one 10-second clip, that is $0.14; Sume 720p models start at $0.75.
- A 12-second Grok Imagine 1.5 Lite clip vs Sume 720p
xAI lists Grok Imagine Video 1.5 Lite at $0.020 a second: $0.24 for one 12-second clip. Sume's grok-imagine-video-1.5 bills $0.15 at 720p.
- A 12-second Vidu Q4 Preview clip vs Sume at 1080p
Vidu Q4 Preview lists launch pricing from $0.014 a second. For one 12-second clip, that is $0.168; Sume 1080p models start at $2.40.
Written by Sume