OmniVoice: 600+ languages, CC-BY-NC weights, hosted TTS instead

OmniVoice covers 600+ languages in a 0.6B model, but its weights are CC-BY-NC. What the card says, what it omits, and where a hosted TTS job fits.

5 min readSume
All posts

OmniVoice is a 0.6B-parameter zero-shot text-to-speech model that its card says covers more than 600 languages, but its pre-trained weights are CC-BY-NC, so they are a research and prototyping tool, not a drop-in engine for paid video voice-over. For a commercial narration pipeline you either train or license something else, or call a hosted TTS job.

Facts below come from the OmniVoice card and its README, read 2026-10-03, and from Sume's docs.

What is OmniVoice, in numbers?

It is listed among the trending text-to-speech models on Hugging Face, with 1.43 million downloads in the last month on the card. The authors describe it as built on a diffusion language model architecture, on top of Qwen3-0.6B, and published the paper on April 1, 2026. It runs on NVIDIA GPUs and Apple Silicon through PyTorch, and installs with pip install omnivoice.

Treat the speed figure as the authors' own number. The card gives a real-time factor but not the GPU, batch size or text length behind it, so a figure of 0.025 tells you the model is fast on their setup, not that your rented card will match it. Before you commit, time ten typical lines on the hardware you will actually pay for, and listen to each one for dropped or mispronounced words.

OmniVoice card and README facts, read 2026-10-03.
ItemWhat the page saysNot stated
LanguagesMore than 600; the card calls it the broadest among zero-shot TTS modelsWhich languages sound good
Voice controlCloning from a short reference, and voice design by gender, age, pitch, dialect or accent, whisperReference length needed
SpeedReal-time factor as low as 0.025, about 40 times faster than real timeWhich GPU
Size0.6B parametersGPU memory
LicenceCode Apache 2.0; pre-trained model CC-BY-NC because of training-data constraintsAny commercial grant

Can you use OmniVoice weights in a paid product?

Not on the licence the card states. The README says the code is Apache 2.0, while the pre-trained model is CC-BY-NC due to constraints from its training data. Non-commercial terms generally rule out selling narration made with the weights or using them inside a paid service. Read the licence file with your counsel before any commercial plan; this post is not legal advice.

The split is the useful lesson: an Apache code repository can sit next to non-commercial weights, and the repository badge tells you nothing about the checkpoint. The same check applies to every model in this series; open-weights licences compared lists the questions to ask.

What does Sume offer for multilingual speech?

Sume's docs name tts_create as a paid hosted-MCP tool (see MCP tools and gates) and say TTS also exists as an HTTP API (Sume basics). The pages I read do not publish a language count for TTS, a voice-design field or a cloning field. Do not assume parity with a 600-language model; check the live catalog and test your target language with one short job before you plan a batch.

Two details help downstream. Avatar videos take an audio_url plus a measured duration_seconds, so any speech file you produce can drive a talking-head job. And video captions take a language hint such as ko or en for the speech-to-text pass, with automatic detection when you omit it.

How do you decide for a low-resource language?

Start from the licence, then the language, then the volume.

  • Non-commercial work, a prototype or a language demo: OmniVoice on a local GPU is a fair way to hear whether your language is covered at all.
  • Paid deliverables: use a hosted job or a model whose weights license commercial use, and keep the evidence of that licence.
  • Rare language with no hosted coverage: ask the hosted catalog first, and if it is missing, budget for a licensed model or a native voice actor instead of leaning on non-commercial weights.
  • Mixed output: generate speech anywhere you are allowed to, then assemble it in Sume with timeline audio and burn captions with a language hint.

Sources

Related posts

More in Models

All Models posts

Written by Sume