OmniVoice: 600+ languages, CC-BY-NC weights, hosted TTS instead
OmniVoice covers 600+ languages in a 0.6B model, but its weights are CC-BY-NC. What the card says, what it omits, and where a hosted TTS job fits.

OmniVoice is a 0.6B-parameter zero-shot text-to-speech model that its card says covers more than 600 languages, but its pre-trained weights are CC-BY-NC, so they are a research and prototyping tool, not a drop-in engine for paid video voice-over. For a commercial narration pipeline you either train or license something else, or call a hosted TTS job.
Facts below come from the OmniVoice card and its README, read 2026-10-03, and from Sume's docs.
What is OmniVoice, in numbers?
It is listed among the trending text-to-speech models on Hugging Face, with 1.43 million downloads in the last month on the card. The authors describe it as built on a diffusion language model architecture, on top of Qwen3-0.6B, and published the paper on April 1, 2026. It runs on NVIDIA GPUs and Apple Silicon through PyTorch, and installs with pip install omnivoice.
Treat the speed figure as the authors' own number. The card gives a real-time factor but not the GPU, batch size or text length behind it, so a figure of 0.025 tells you the model is fast on their setup, not that your rented card will match it. Before you commit, time ten typical lines on the hardware you will actually pay for, and listen to each one for dropped or mispronounced words.
| Item | What the page says | Not stated |
|---|---|---|
| Languages | More than 600; the card calls it the broadest among zero-shot TTS models | Which languages sound good |
| Voice control | Cloning from a short reference, and voice design by gender, age, pitch, dialect or accent, whisper | Reference length needed |
| Speed | Real-time factor as low as 0.025, about 40 times faster than real time | Which GPU |
| Size | 0.6B parameters | GPU memory |
| Licence | Code Apache 2.0; pre-trained model CC-BY-NC because of training-data constraints | Any commercial grant |
Can you use OmniVoice weights in a paid product?
Not on the licence the card states. The README says the code is Apache 2.0, while the pre-trained model is CC-BY-NC due to constraints from its training data. Non-commercial terms generally rule out selling narration made with the weights or using them inside a paid service. Read the licence file with your counsel before any commercial plan; this post is not legal advice.
The split is the useful lesson: an Apache code repository can sit next to non-commercial weights, and the repository badge tells you nothing about the checkpoint. The same check applies to every model in this series; open-weights licences compared lists the questions to ask.
What does Sume offer for multilingual speech?
Sume's docs name tts_create as a paid hosted-MCP tool (see MCP tools and gates) and say TTS also exists as an HTTP API (Sume basics). The pages I read do not publish a language count for TTS, a voice-design field or a cloning field. Do not assume parity with a 600-language model; check the live catalog and test your target language with one short job before you plan a batch.
Two details help downstream. Avatar videos take an audio_url plus a measured duration_seconds, so any speech file you produce can drive a talking-head job. And video captions take a language hint such as ko or en for the speech-to-text pass, with automatic detection when you omit it.
How do you decide for a low-resource language?
Start from the licence, then the language, then the volume.
- Non-commercial work, a prototype or a language demo: OmniVoice on a local GPU is a fair way to hear whether your language is covered at all.
- Paid deliverables: use a hosted job or a model whose weights license commercial use, and keep the evidence of that licence.
- Rare language with no hosted coverage: ask the hosted catalog first, and if it is missing, budget for a licensed model or a native voice actor instead of leaning on non-commercial weights.
- Mixed output: generate speech anywhere you are allowed to, then assemble it in Sume with timeline audio and burn captions with a language hint.
Sources
Related posts
More in Models
- Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.
- Polish, Dutch, Swedish, Turkish text to speech API: Sume pl nl sv tr
Sume's Voices library tags voices pl, nl, sv and tr alongside 12 other languages. What to send for each, and how Eleven v4's list compares.
- QuantFunc INT4 MiniMax H3: 3.2 s per step on an RTX 4090
QuantFunc's 4-bit MiniMax H3 claims 3.2 s per step on an RTX 4090. That is not a clip time. What the card says, what it omits, and when to use a hosted job.
- Reference audio for AI video: which models accept a voice clip
Seedance 2.5, Wan 3.0 and MiniMax H3 take reference audio on Sume; Kling 3, Gemini Omni Flash and the swap rows do not. Limits and the one-reference rule.
Written by Sume