Mochi 1 LoRA on one H100 vs reference images on a hosted API
Mochi 1's trainer fine-tunes a LoRA on one H100 or A100 80GB. Sume has no LoRA field, so here is when to train and when to send reference images.

Mochi 1's repository ships a LoRA trainer that, per the README, can fine-tune the model on one H100 or A100 80GB GPU. Sume's video request has no LoRA, checkpoint or training field, so if a custom-trained style is the goal you train Mochi yourself; if consistency from a few pictures is the goal, reference images on a hosted model are the faster route.
What is Mochi 1?
The README describes a 10B-parameter AsymmDiT diffusion transformer with 48 layers and 24 attention heads, released under Apache 2.0. It generates 480p video, 31 frames per generation, and is optimized for photorealistic content; the project says it does not perform well with animated content, and that extreme motion can cause warping.
On memory, the README says about 60GB VRAM on a single GPU for the main repository, with a ComfyUI path that reduces this to under 20GB. ComfyUI consumer-GPU support was added on November 5, 2024, and the LoRA trainer on November 26, 2024.
What does a LoRA give you that references do not?
A trained LoRA bakes a subject or style into the weights, so it applies to any prompt without re-supplying pictures. It costs you a dataset of clips, GPU time on an 80GB card and a training run. References, where a model supports them, steer a single generation with no training.
| Need | Mochi 1 LoRA | Hosted reference images |
|---|---|---|
| Own style on every prompt | Yes, after training | Resend references each time |
| Setup | H100 or A100 80GB, dataset | None |
| Resolution | 480p | Per model, see catalog |
| Weights stay private | Yes | No, media goes to the provider |
| Works with animated content | Weak per README | Depends on model |
What does Sume accept instead?
POST /v1/videos takes frame_images for first and last frames and input_references for reference-to-video, but only for models whose supported_input_references lists the type. wan-3.0, for example, honors image, video and audio references per the docs. The catalog from GET /v1/videos/models is the only reliable list of what each id accepts.
The Video generation docs do not list Mochi as a catalog id, so I would not plan on calling it hosted. See Mochi's VRAM and limits for the model itself.
How do I decide?
Train a LoRA when the look must be exactly yours across hundreds of clips and you own the GPU. Use references when you have a handful of images and a one-off deadline. If you try references first and the result is close enough, you have saved a training run; if it is not, you have a clear argument for training.
What would a practical test look like?
Pick the look you want and gather five to ten stills. Send them as input_references to a hosted model that lists the type in supported_input_references, and generate three short clips at the lowest resolution the model offers. If the subject holds across the three, you probably do not need to train anything.
If it drifts, write down how: face, colour, outfit or motion. A LoRA is most useful for drift you can describe and reproduce in a dataset. Mochi's own README warns about warping with extreme motion and weak results on animated content, so check those limits against your footage before you rent an 80GB GPU.
Also budget for the safety note in the README: the project says commercial deployment needs additional safety protocols, which is your work, not the weights'.
Sources
Related posts
More in Models
- Where are the lyrics in a Lyria 3.5 result? Gemini vs Sume
Google returns Lyria 3.5 lyrics and song structure as text beside the audio. Sume puts model-reported lyrics or a section map in result.lyrics when present.
- Nano Banana 2 Lite model id: Vertex hyphens vs Gemini API dots
Google's Gemini API names Nano Banana 2 Lite gemini-3.1-flash-lite-image; its Cloud post writes gemini-3-1-flash-lite-image. Sume's image catalog lists neither.
- Nano Banana negative prompt: describe what you want instead
Google's Gemini image guide says to write semantic negative prompts: describe an empty street, not 'no cars'. Sume's image request has no negative field.
- Nano Banana prompt languages: Korean and Japanese, per Google
Google lists the languages Gemini image models work best in, including ko-KR and ja-JP. What that means for a Korean or Japanese prompt sent through Sume.
Written by Sume