Mochi video model: what Mochi 1 is, needs and can't do

Mochi 1 preview is Genmo's Apache 2.0 open video model: 480p output, about 60 GB VRAM on one GPU, and it is not in Sume's video catalog.

4 min readSume
All posts

Mochi 1 is an open text-to-video model from Genmo. Its Hugging Face card calls the current release "Mochi 1 preview", says it is released under the Apache 2.0 license, and says the initial release generates video at 480p. It is downloadable weights, and Sume's video catalog does not include it.

The Mochi facts are from Genmo's model card, read 2026-09-29. The Sume facts are from Video generation.

How much VRAM does Mochi 1 need?

It depends on which code path you use. The card gives three figures, each tied to a specific way of running the model.

From the Mochi 1 preview model card, read 2026-09-29.
Way of running itVRAM the card states
Genmo's repository, single GPUApproximately 60GB; the card recommends at least 1 H100 GPU
Diffusers, highest quality example42GB
Diffusers, bfloat16 variant22GB, with a slight drop in quality

How big is Mochi 1?

The card describes Mochi 1 as a 10 billion parameter diffusion model with a novel asymmetric architecture. Its AsymmDiT processes the prompt alongside compressed video tokens, and the card says Mochi 1 encodes prompts with a single T5-XXL language model instead of several pretrained language models. It also says it expects the community to fine-tune the model to suit various aesthetic preferences.

What does Mochi 1 not do well?

The card's Limitations section is direct about the preview status. It says Mochi 1 is "a living and evolving checkpoint" and lists these known limits:

  • Video is generated at 480p today.
  • In some edge cases with extreme motion, minor warping and distortions can occur.
  • It is optimized for photorealistic styles and does not perform well with animated content.

How do I run Mochi 1?

The card installs Genmo's repository with uv, tells you to install FFMPEG to turn outputs into videos, and offers two entry points: a Gradio UI and a command-line script.

  • Gradio UI: python3 ./demos/gradio_ui.py --model_dir "<path_to_downloaded_directory>".
  • CLI: python3 ./demos/cli.py --model_dir "<path_to_downloaded_directory>".
  • The card says ComfyUI can optimize Mochi to run on less than 20GB VRAM, while Genmo's own implementation prioritizes flexibility over memory efficiency.

Can I use Mochi 1 commercially?

The card states the Apache 2.0 license. Its Safety section adds that organizations should implement additional safety protocols and careful consideration before deploying the weights in any commercial services or products. Treat that as the model authors' guidance, and read the license itself for terms.

Is Mochi available through the Sume API?

No. The video ids in Sume's catalog code are seedance-2.5, seedance-2-mini, seedance-2, seedance-2-fast, kling-3, wan-3.0, grok-imagine-video-1.5, minimax-h3, minimax-h3-max, gemini-omni-flash-1.1; Mochi is not one of them, and the video generation docs do not mention it. You would run the weights on your own GPU, or pick a listed model.

Open-source video model vs API compares the two routes, and Do you need a GPU for AI video? covers the hardware side.

Sources

Related posts

More in Models

All Models posts

Written by Sume