HunyuanVideo 1.5 LoRA training: train.py and the Muon optimizer
HunyuanVideo 1.5 released training code on Dec 5 2025 and says to use the Muon optimizer for LoRA. The flags, the torchrun and FSDP setup, and the hosted gap.

Yes, HunyuanVideo 1.5 has official LoRA training code. The README's Dec 05, 2025 news line says training code and a LoRA tuning script were released with Muon optimizer support, and the README states that to continue training the model or fine-tune it with LoRA you should use the Muon optimizer. The entry point is train.py, with --use_lora, --lora_r 8 and --lora_alpha 16 in the example.
What does the training setup need?
From the README: distributed training launched with torchrun, FSDP support, and gradient checkpointing. It names Muon as the optimizer to use for continued training and LoRA. The README gives no VRAM figure for training, and I did not find a dataset-size or training-time example, so budget by testing a short run first.
| Piece | What the README says |
|---|---|
| Script | train.py |
| Optimizer | Muon |
| LoRA flags | --use_lora, --lora_r 8, --lora_alpha 16 |
| Launch | torchrun |
| Memory features | FSDP, gradient checkpointing |
| Released | Dec 05, 2025 |
What does a hosted API offer instead?
Not a LoRA. In v1, Sume's allowed_passthrough_parameters list is empty for every video model and a non-empty provider.options returns 400 unsupported_parameter, so there is no field to attach your own weights to. The route for a look or a character on Sume is reference media: input_references for style or content, or frame_images for a first or last frame, on the models that list them.
So the choice is concrete. If your product depends on a trained style that must be applied to every clip, you need open weights and your own GPUs, and the license that comes with them. If a few reference images are enough, a hosted id is less work. Sume does not list HunyuanVideo (catalog code, read 2026-10-09); see the video docs for what is listed.
What should a first training run look like?
Start small: a short clip set of one subject or one style, a low rank as in the README example (--lora_r 8), and a few checkpoints saved early. Test each checkpoint on prompts that are not in the training set, to see whether you have learned the look or memorized the clips. The README does not give recommended dataset sizes or epochs.
Before you train on anyone's footage, confirm you have the right to use it. Training data rights are your responsibility, and check the base weights' license for what it says about derivatives.
Sources
Related posts
More in Developers
- HunyuanVideo 1.5 cache inference: DeepCache, TeaCache, TaylorCache
HunyuanVideo 1.5 added DeepCache on Nov 24 2025 and TeaCache plus TaylorCache on Nov 27, switched with --enable_cache and --cache_type. What the README claims.
- Hy Image 3.5 multi-turn editing with assembled_history vs Sume edits
How Tencent's Hy Image 3.5 Preview chains edits with assembled_history, and what the same loop looks like on Sume's stateless images route.
- Idempotency key from the order id, not a fresh UUID per attempt
A random UUID generated inside the retry loop gives every attempt a new key and a new paid job. Derive the Sume Idempotency-Key from the order instead.
- Inspect transcript words without start or end: skip them in cuts
Inspect transcript words can omit start or end. Treat an untimed word as unknown, and never cut across a gap that contains one: 0 of its time is safe to remove.
Written by Sume