Fine-tune MiniMax H3 with a LoRA: data grid, recipe and serving

MiniMax points to an experimental miles-diffusion LoRA recipe for H3: 1344x768, 24 fps, 17n+5 frames, 8 x H200. The data rules and how to serve it.

5 min readSume
All posts

MiniMax's self-hosting guide links an experimental LoRA fine-tuning recipe for H3, built on the miles-diffusion repository and reviewed upstream on August 26, 2026. Your clips must sit on a strict grid: 1344 x 768 (768 short edge, 16:9), 24 fps within 0.01, and a frame count of 17n + 5, with a minimum of 107 frames, about 4.46 seconds.

All figures here come from the Run and Self-host MiniMax H3 guide, read on 2026-10-02. The guide calls the recipe experimental, so test before you build on it. A LoRA you train runs on your own SGLang server, not on a hosted API such as Sume's.

What must the training data look like?

Each record is a JSON line with a prompt and a metadata object that points at a clip, for example {"prompt": "...", "metadata": {"video": "clips/clip.mp4"}}. The frame rule is the part that bites. At 24 fps, 17n + 5 frames means 107, 124, 141 and so on, so a 5-second clip of 120 frames is off the grid and needs trimming or re-encoding.

A quick calculation helps: frames = 17n + 5 gives 107 at n = 6 and 124 at n = 7, which is the 124-frame length the guide uses for its latency figure.

Training data rules from MiniMax, read 2026-10-02
RuleValue
Canvas768 short edge, 16:9, so 1344 x 768
Frame rate24 fps, strict to 0.01
Frames17n + 5, minimum 107 (about 4.46 s)
RecordJSONL with prompt and metadata.video

What is the recipe and what does it cost?

The defaults are a learning rate of 3e-5, which the guide calls the most sensitive setting, and LoRA rank 64 with alpha 128. The reference hardware is 8 x H200. For 10 epochs the guide reports about 66 minutes on a cold cache and about 2 hours 31 minutes on a warm cache; the default is 3 epochs. It does not state a dollar cost.

The training script writes checkpoints, and a separate script exports a LoRA file with the same rank and alpha.

python3 scripts/run_diffusion_sft_h3_t2va.py
# custom data
#   --extra-args "--prompt-data /abs/train.jsonl"
# longer run
#   --num-epoch 10
python3 scripts/export_lora.py \
  --ckpt-dir <run>/ckpt/iter_0000070 \
  --out my_h3_lora.safetensors \
  --lora-rank 64 --lora-alpha 128

How do you serve the result and what are the limits?

In SGLang you pass --lora-path my_h3_lora.safetensors and keep the adapter_config.json sidecar file next to it. The recipe targets the text-to-video path; Ref2VA is a separate checkpoint, as covered in our FL2VA vs Ref2VA post.

The license applies to what you build. The guide names the MiniMax H3 Community License Agreement and lists the US, EU, UK and South Korea as Excluded Territories requiring a separate license. Read it before training on, or shipping, a derived model.

Should you fine-tune or use a hosted model?

Fine-tune when you need one look that prompts and references cannot hold, and you can run 8 H200-class GPUs. Otherwise, try references first: Sume's minimax-h3 accepts reference images, video and audio on a hosted request, with limits listed per model in supported_input_references. A reference is a hint on one request; a LoRA changes the model. Which one you need depends on whether the look must survive across thousands of prompts.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume