Mochi 1 LoRA on one H100 vs reference images on a hosted API

Mochi 1's trainer fine-tunes a LoRA on one H100 or A100 80GB. Sume has no LoRA field, so here is when to train and when to send reference images.

4 min readSume
All posts

Mochi 1's repository ships a LoRA trainer that, per the README, can fine-tune the model on one H100 or A100 80GB GPU. Sume's video request has no LoRA, checkpoint or training field, so if a custom-trained style is the goal you train Mochi yourself; if consistency from a few pictures is the goal, reference images on a hosted model are the faster route.

What is Mochi 1?

The README describes a 10B-parameter AsymmDiT diffusion transformer with 48 layers and 24 attention heads, released under Apache 2.0. It generates 480p video, 31 frames per generation, and is optimized for photorealistic content; the project says it does not perform well with animated content, and that extreme motion can cause warping.

On memory, the README says about 60GB VRAM on a single GPU for the main repository, with a ComfyUI path that reduces this to under 20GB. ComfyUI consumer-GPU support was added on November 5, 2024, and the LoRA trainer on November 26, 2024.

What does a LoRA give you that references do not?

A trained LoRA bakes a subject or style into the weights, so it applies to any prompt without re-supplying pictures. It costs you a dataset of clips, GPU time on an 80GB card and a training run. References, where a model supports them, steer a single generation with no training.

Fine-tune vs reference images (read 2026-10-02)
NeedMochi 1 LoRAHosted reference images
Own style on every promptYes, after trainingResend references each time
SetupH100 or A100 80GB, datasetNone
Resolution480pPer model, see catalog
Weights stay privateYesNo, media goes to the provider
Works with animated contentWeak per READMEDepends on model

What does Sume accept instead?

POST /v1/videos takes frame_images for first and last frames and input_references for reference-to-video, but only for models whose supported_input_references lists the type. wan-3.0, for example, honors image, video and audio references per the docs. The catalog from GET /v1/videos/models is the only reliable list of what each id accepts.

The Video generation docs do not list Mochi as a catalog id, so I would not plan on calling it hosted. See Mochi's VRAM and limits for the model itself.

How do I decide?

Train a LoRA when the look must be exactly yours across hundreds of clips and you own the GPU. Use references when you have a handful of images and a one-off deadline. If you try references first and the result is close enough, you have saved a training run; if it is not, you have a clear argument for training.

What would a practical test look like?

Pick the look you want and gather five to ten stills. Send them as input_references to a hosted model that lists the type in supported_input_references, and generate three short clips at the lowest resolution the model offers. If the subject holds across the three, you probably do not need to train anything.

If it drifts, write down how: face, colour, outfit or motion. A LoRA is most useful for drift you can describe and reproduce in a dataset. Mochi's own README warns about warping with extreme motion and weak results on animated content, so check those limits against your footage before you rent an 80GB GPU.

Also budget for the safety note in the README: the project says commercial deployment needs additional safety protocols, which is your work, not the weights'.

Sources

Related posts

More in Models

All Models posts

Written by Sume