Qwen3-TTS (Apache 2.0): 3-second clones and 97 ms vs Sume TTS
Qwen3-TTS is Apache 2.0 with 3-second voice cloning, natural-language voice design and 97 ms streaming. What Sume's hosted TTS route does and does not cover.

Qwen3-TTS is an Apache 2.0 open model family from the Qwen team that clones a voice from about 3 seconds of audio and can design a voice from a text description, and you run it yourself. Sume's TTS route does neither of those at request time: you pick a Cartesia Sonic engine and send a voice id, and the job is billed per character.
Qwen3-TTS facts below come from the QwenLM repository, read on 2026-10-02. Sume facts come from the Sume API reference.
What does the Qwen3-TTS repository list?
The README announces the release on 2026-01-22 under the Apache-2.0 license. Supported languages are Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. Model sizes are Qwen3-TTS-12Hz-0.6B (two variants) and Qwen3-TTS-12Hz-1.7B (three variants).
The base models support a 3-second rapid voice clone from user audio, and the README says both the reference audio and its transcript are needed for best quality. Voice design, which generates speech from natural-language instructions about timbre, emotion and prosody, is available only in the 1.7B-VoiceDesign model. Streaming end-to-end latency is stated as low as 97 ms, with the first audio packet emitted after a single character of input.
Which Sume features line up with that?
Sume's controls are narrower and more explicit. generation_config takes speed from 0.6 to 1.5, volume from 0.5 to 2.0, and a free-text emotion guide of up to 64 characters. language is a BCP-47 or ISO-639 code that you should set for every non-English transcript. There is no field for describing a new voice, and the public contract lists no route that creates one, so voices are chosen, not designed, per request.
| Capability | Qwen3-TTS | Sume TTS Router |
|---|---|---|
| Voice cloning | 3-second reference plus transcript, run locally | Not a request field; use an existing voice id |
| Voice design from text | Yes, in the 1.7B-VoiceDesign model | No |
| Latency target | As low as 97 ms streaming, per the README | Asynchronous job; poll status and result |
| Languages | 10 listed in the README | Set the language field; a known voice-language mismatch returns 409 before charge |
| License | Apache-2.0 | Hosted, billed at $0.0475 per 1,000 characters |
Which workload belongs where?
Streaming at 97 ms is built for live voice agents and interactive playback. A job queue is built for rendered media: narration for a video, a podcast intro, a batch of product lines. If you need sub-second playback, an asynchronous HTTP job is the wrong tool and the Qwen3-TTS streaming mode is a closer match.
If you are producing a finished file, the shape of the Sume route is an advantage: one request, one job id, a Sume-hosted URL, optional word timestamps and sentence slices. Words and sentence timings let you cut captions or slide changes from the audio you generated.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare two ways to get a voiceover.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en"
}'
Then fetch the job:
curl https://api.sume.com/v1/jobs/$JOB_ID/status \
-H "Authorization: Bearer $SUME_API_KEY"
curl https://api.sume.com/v1/jobs/$JOB_ID/result \
-H "Authorization: Bearer $SUME_API_KEY"
What should you check before choosing?
Pick by constraint, not by headline number.
- Do you need to clone a new voice per request? Then a local model is the only option here.
- Do you need playback while text is still arriving? Then streaming matters more than anything else.
- Do you need a billed, repeatable job with an id and a hosted file? Then a hosted route removes the operations work.
- Do you have permission to clone each voice you use? Consent is required on any engine, and synthetic speech should be labeled where rules ask for it.
Can the two coexist?
Yes. Teams often prototype with open weights, measure what a script sounds like, then move final renders to a hosted voice for repeatability. The cost of switching is mostly pronunciation tuning and voice choice, so keep your test lines and compare them side by side. Use the same transcript on both and listen for names, numbers and non-English words before you commit.
One more difference is who carries the failure modes. A local model fails when your GPU runs out of memory or a dependency changes. A hosted job fails with a stable error code and, for a refused request, before credit is reserved. Neither is better in general; they are different bargains, and the right one depends on whether your team would rather tune inference or write the rest of the product.
Sources
Related posts
More in Comparisons
- Realtime AI video restyle: Decart Lucy over WebRTC vs Sume jobs
Decart's realtime Lucy models restyle a live feed over WebRTC at 25 FPS and 1280x704. Sume's video edit is an async job on a finished clip, not a live stream.
- Recraft API has 15 endpoints: which does the Sume image API cover?
Recraft lists 15 API endpoints, from vectorize to erase region. Sume's image API covers generate and reference edit with recraft/recraft-v4, not the rest.
- Replicate output files deleted after an hour: what Sume's URLs do
On Replicate, files from API predictions are deleted after an hour. Sume's Format run docs call media URLs durable, but check expires_at when a URL is signed.
- Replicate deletes API outputs after an hour: what Sume keeps
Replicate removes API prediction inputs, outputs and files after an hour by default. Sume result URLs on media.sume.com do not expire. What to save and when.
Written by Sume