Irodori-TTS-v4-Large: Japanese cloning, 120 s reference, terms

Irodori-TTS-v4-Large is a 3.29B Japanese TTS model with emoji style control and Gemma terms. What the card says and how Sume's audio tools fit.

5 min readSume
All posts

Irodori-TTS-v4-Large is a Japanese-only, 3.29B-parameter text-to-speech model that clones a voice from up to 120 seconds of reference audio and lets you set the style with emoji; its weights fall under the Gemma Terms of Use, not a plain open licence. Use it when Japanese voice quality is the whole job and you can run a large model. For mixed-language batches or a no-GPU pipeline, a hosted TTS job and Sume's assembly tools are simpler.

Model facts here come from the Irodori-TTS-v4-Large card, read 2026-10-03, where it appears among Hugging Face's trending text-to-speech models.

What does the model card say about Irodori-TTS-v4-Large?

The card describes a Rectified Flow Diffusion Transformer with 24 layers at 2,048 dimensions, 3.29 billion parameters in total. It supports zero-shot voice cloning, text-based voice design and style-controlled cloning, with emoji as the style control. The language is Japanese only.

The card reports two speaker-similarity scores, measured as CAM++ cosine similarity: 0.7593 with 30 seconds of reference audio and 0.7788 with 120 seconds, the longest combined reference it accepts. That is a small gain for four times the reference audio, which is useful when you decide how much of a speaker's recording to collect.

Two things the card does not say are worth knowing before you plan around it: it gives no GPU memory requirement and no maximum output length, and the quantized builds that would shrink it are listed as planned. Budget a test day on your own hardware rather than a purchase decision from the page alone.

Irodori-TTS-v4-Large card facts, read 2026-10-03.
ItemCard value
LanguageJapanese only
ArchitectureRectified Flow Diffusion Transformer, 24 layers, 2,048 dimensions
Parameters3.29 billion
Reference audio for cloningUp to 120 seconds combined
Speaker similarity (CAM++ cosine)0.7593 at 30 s; 0.7788 at 120 s
Quantized variantsINT8, INT4 and FP8 are planned, not released
TermsGemma Terms of Use, through a text encoder derived from google/t5gemma-2-1b-1b

What does the Gemma licence mean for a commercial voice?

The card says the model is subject to the Gemma Terms of Use because its text encoder derives from a Google T5Gemma checkpoint. That is a use-based licence with its own acceptable-use rules, which is different from Apache 2.0 or MIT. Read the terms yourself before shipping paid audio, and check whether they pass to anything fine-tuned from this model.

Separately, cloning a real person's voice needs that person's consent, whatever the licence says. The card's similarity scores measure how close the clone gets, which is also how close it gets to impersonation.

What can Sume do around a Japanese voice?

Sume's docs name tts_create as a paid tool in the hosted MCP inventory. The pages I read do not list a Japanese-specific voice, a cloning field or an emoji style control, so test a short Japanese line in the live catalog before committing, and do not assume Irodori-level cloning.

What Sume does handle is everything after the audio exists. Avatar videos accept an audio_url and a measured duration_seconds, so speech made elsewhere can drive a talking head. Timeline audio joins up to 20 audio parts without re-synthesis. Video captions take a language code for the transcription pass, such as ko or en, and detect automatically when you omit it.

Run it yourself or call a hosted route?

The deciding facts are the size and the quantization. A 3.29B model needs a real GPU, and the card says the smaller INT8, INT4 and FP8 builds are still planned, so you cannot plan on a small footprint today.

If you do run it, keep a short test script: one line per style emoji, one clone from 30 seconds and one from 120 seconds, so you can hear whether the extra reference is worth collecting for your speaker.

  • Japanese-only, high-fidelity cloning with consent and a GPU: Irodori-TTS-v4-Large is a candidate.
  • Several languages in one batch: a single hosted catalog is easier than one local model per language.
  • Nothing to host: call a hosted job and assemble with Sume's audio tools.
  • Unsure of terms: read the Gemma Terms of Use before any paid use.

Sources

Related posts

More in Models

All Models posts

Written by Sume