Qwen3-TTS (Apache 2.0): 3-second clones and 97 ms vs Sume TTS

Qwen3-TTS is Apache 2.0 with 3-second voice cloning, natural-language voice design and 97 ms streaming. What Sume's hosted TTS route does and does not cover.

5 min readSume
All posts

Qwen3-TTS is an Apache 2.0 open model family from the Qwen team that clones a voice from about 3 seconds of audio and can design a voice from a text description, and you run it yourself. Sume's TTS route does neither of those at request time: you pick a Cartesia Sonic engine and send a voice id, and the job is billed per character.

Qwen3-TTS facts below come from the QwenLM repository, read on 2026-10-02. Sume facts come from the Sume API reference.

What does the Qwen3-TTS repository list?

The README announces the release on 2026-01-22 under the Apache-2.0 license. Supported languages are Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. Model sizes are Qwen3-TTS-12Hz-0.6B (two variants) and Qwen3-TTS-12Hz-1.7B (three variants).

The base models support a 3-second rapid voice clone from user audio, and the README says both the reference audio and its transcript are needed for best quality. Voice design, which generates speech from natural-language instructions about timbre, emotion and prosody, is available only in the 1.7B-VoiceDesign model. Streaming end-to-end latency is stated as low as 97 ms, with the first audio packet emitted after a single character of input.

Which Sume features line up with that?

Sume's controls are narrower and more explicit. generation_config takes speed from 0.6 to 1.5, volume from 0.5 to 2.0, and a free-text emotion guide of up to 64 characters. language is a BCP-47 or ISO-639 code that you should set for every non-English transcript. There is no field for describing a new voice, and the public contract lists no route that creates one, so voices are chosen, not designed, per request.

Qwen3-TTS versus Sume TTS Router, read 2026-10-02
CapabilityQwen3-TTSSume TTS Router
Voice cloning3-second reference plus transcript, run locallyNot a request field; use an existing voice id
Voice design from textYes, in the 1.7B-VoiceDesign modelNo
Latency targetAs low as 97 ms streaming, per the READMEAsynchronous job; poll status and result
Languages10 listed in the READMESet the language field; a known voice-language mismatch returns 409 before charge
LicenseApache-2.0Hosted, billed at $0.0475 per 1,000 characters

Which workload belongs where?

Streaming at 97 ms is built for live voice agents and interactive playback. A job queue is built for rendered media: narration for a video, a podcast intro, a batch of product lines. If you need sub-second playback, an asynchronous HTTP job is the wrong tool and the Qwen3-TTS streaming mode is a closer match.

If you are producing a finished file, the shape of the Sume route is an advantage: one request, one job id, a Sume-hosted URL, optional word timestamps and sentence slices. Words and sentence timings let you cut captions or slide changes from the audio you generated.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: narration-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Welcome back. Today we compare two ways to get a voiceover.",
    "voice": { "id": "'"$VOICE_ID"'" },
    "language": "en"
  }'

Then fetch the job:

curl https://api.sume.com/v1/jobs/$JOB_ID/status \
  -H "Authorization: Bearer $SUME_API_KEY"

curl https://api.sume.com/v1/jobs/$JOB_ID/result \
  -H "Authorization: Bearer $SUME_API_KEY"

What should you check before choosing?

Pick by constraint, not by headline number.

  • Do you need to clone a new voice per request? Then a local model is the only option here.
  • Do you need playback while text is still arriving? Then streaming matters more than anything else.
  • Do you need a billed, repeatable job with an id and a hosted file? Then a hosted route removes the operations work.
  • Do you have permission to clone each voice you use? Consent is required on any engine, and synthetic speech should be labeled where rules ask for it.

Can the two coexist?

Yes. Teams often prototype with open weights, measure what a script sounds like, then move final renders to a hosted voice for repeatability. The cost of switching is mostly pronunciation tuning and voice choice, so keep your test lines and compare them side by side. Use the same transcript on both and listen for names, numbers and non-English words before you commit.

One more difference is who carries the failure modes. A local model fails when your GPU runs out of memory or a dependency changes. A hosted job fails with a stable error code and, for a refused request, before credit is reserved. Neither is better in general; they are different bargains, and the right one depends on whether your team would rather tune inference or write the rest of the product.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume