NVIDIA Magpie TTS for voice agents vs Sume TTS 1.0: which to use?

Magpie TTS is an open-weights, low-latency model for live voice agents in 12 languages. Sume TTS is async and non-streaming. How to choose, read 2026-10-10.

4 min readSume
All posts

For a live voice agent that must start speaking in tens of milliseconds, Sume TTS 1.0 is the wrong tool and NVIDIA Magpie TTS is a better fit; for narration, video voiceover and any audio you render ahead of time, Sume TTS is the simpler one. Sume TTS 1.0 is async-only and non-streaming, so it returns a finished file rather than audio you can play while it is still being made.

The Magpie facts below come from NVIDIA's Hugging Face blog post of August 10, 2026, read on 2026-10-10. Sume facts come from the repository and docs.

What NVIDIA says about Magpie TTS

The post describes a 364-million-parameter open-weights text-to-speech model under the NVIDIA Open Model License, covering 12 languages: English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, Modern Standard Arabic, Korean and Brazilian Portuguese. It offers a male and a female voice per language, and code-switching for Hindi and Japanese. The post does not mention voice cloning or SSML.

The headline numbers are time to first audio: 32 ms on a B200, 47 ms on an H100, 53 ms on a DGX Spark and 79 ms on an A100. With 64 concurrent streams on a B200 the figure rises to 239 ms. These are NVIDIA's measurements on NVIDIA hardware, not something Sume has reproduced. You deploy it through NIM or the open checkpoint, on GPUs you run.

Magpie TTS time to first audio as reported by NVIDIA, read 2026-10-10
HardwareTime to first audioLoad
B20032 msSingle stream
H10047 msSingle stream
DGX Spark53 msSingle stream
A10079 msSingle stream
B200239 ms64 concurrent streams

What Sume TTS 1.0 does

Sume TTS 1.0 takes a transcript of up to 20,000 characters and a voice, and returns an audio file through a job. Modes are async, sync, subscribe and webhook; sync and subscribe wait at most 30 seconds before you poll. There is no audio stream, so there is no time-to-first-audio number to compare, and this post does not invent one.

Pricing is $0.0475 per 1,000 characters. Synthesized audio over 1,200 seconds fails with tts_duration_exceeded, so long scripts are split. Output formats are mp3, wav or raw, with sample rates from 8,000 to 48,000 Hz, and word timestamps are available when you need captions or edit points. Voices come from avatar_id, avatar_handle or a voice id; voice cloning is app-only and not in the API.

Pick by job, not by model

The deciding fact is whether a person is waiting on the other end of a live conversation. If yes, you need a streaming engine whose first audio arrives almost immediately, and open weights on your own GPUs is a legitimate way to get it. If the audio is rendered, reviewed and then published, async is fine and you avoid running anything.

  • Live phone or in-app agent: choose a streaming model such as Magpie, and budget GPUs, scaling and monitoring yourself.
  • Video voiceover, explainer or ad read: choose Sume TTS and join takes with Timeline audio when you need one file.
  • Compliance needs on-premises audio: open weights keep text on your network; Sume is a hosted API.
  • Language outside Magpie's 12 or a specific accent: check each engine's voice list directly before committing.

A hybrid that keeps both honest

Many teams need both: a live agent for conversation and prerecorded audio for the content around it, such as a welcome message, a hold announcement or a product video. Nothing stops you from using a streaming model for the first and Sume for the second, as long as the voices are close enough that callers do not notice the handoff.

To see how Sume jobs are polled and when a result is ready, read jobs and results. For the latency question in the abstract, 150 ms claims vs async TTS goes through which jobs need which. For another open-weights comparison see Voxtral TTS.

Limits of this comparison

It does not compare naturalness, because neither source supports a ranking. It also does not price the self-hosted route: GPU hours, ops time and licence review sit with you. Re-read NVIDIA's license terms before shipping, since the Open Model License has conditions that this post does not summarize.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume