Fine-tune MiniMax H3 with a LoRA: data grid, recipe and serving
MiniMax points to an experimental miles-diffusion LoRA recipe for H3: 1344x768, 24 fps, 17n+5 frames, 8 x H200. The data rules and how to serve it.

MiniMax's self-hosting guide links an experimental LoRA fine-tuning recipe for H3, built on the miles-diffusion repository and reviewed upstream on August 26, 2026. Your clips must sit on a strict grid: 1344 x 768 (768 short edge, 16:9), 24 fps within 0.01, and a frame count of 17n + 5, with a minimum of 107 frames, about 4.46 seconds.
All figures here come from the Run and Self-host MiniMax H3 guide, read on 2026-10-02. The guide calls the recipe experimental, so test before you build on it. A LoRA you train runs on your own SGLang server, not on a hosted API such as Sume's.
What must the training data look like?
Each record is a JSON line with a prompt and a metadata object that points at a clip, for example {"prompt": "...", "metadata": {"video": "clips/clip.mp4"}}. The frame rule is the part that bites. At 24 fps, 17n + 5 frames means 107, 124, 141 and so on, so a 5-second clip of 120 frames is off the grid and needs trimming or re-encoding.
A quick calculation helps: frames = 17n + 5 gives 107 at n = 6 and 124 at n = 7, which is the 124-frame length the guide uses for its latency figure.
| Rule | Value |
|---|---|
| Canvas | 768 short edge, 16:9, so 1344 x 768 |
| Frame rate | 24 fps, strict to 0.01 |
| Frames | 17n + 5, minimum 107 (about 4.46 s) |
| Record | JSONL with prompt and metadata.video |
What is the recipe and what does it cost?
The defaults are a learning rate of 3e-5, which the guide calls the most sensitive setting, and LoRA rank 64 with alpha 128. The reference hardware is 8 x H200. For 10 epochs the guide reports about 66 minutes on a cold cache and about 2 hours 31 minutes on a warm cache; the default is 3 epochs. It does not state a dollar cost.
The training script writes checkpoints, and a separate script exports a LoRA file with the same rank and alpha.
python3 scripts/run_diffusion_sft_h3_t2va.py
# custom data
# --extra-args "--prompt-data /abs/train.jsonl"
# longer run
# --num-epoch 10
python3 scripts/export_lora.py \
--ckpt-dir <run>/ckpt/iter_0000070 \
--out my_h3_lora.safetensors \
--lora-rank 64 --lora-alpha 128How do you serve the result and what are the limits?
In SGLang you pass --lora-path my_h3_lora.safetensors and keep the adapter_config.json sidecar file next to it. The recipe targets the text-to-video path; Ref2VA is a separate checkpoint, as covered in our FL2VA vs Ref2VA post.
The license applies to what you build. The guide names the MiniMax H3 Community License Agreement and lists the US, EU, UK and South Korea as Excluded Territories requiring a separate license. Read it before training on, or shipping, a derived model.
Should you fine-tune or use a hosted model?
Fine-tune when you need one look that prompts and references cannot hold, and you can run 8 H200-class GPUs. Otherwise, try references first: Sume's minimax-h3 accepts reference images, video and audio on a hosted request, with limits listed per model in supported_input_references. A reference is a hint on one request; a LoRA changes the model. Which one you need depends on whether the look must survive across thousands of prompts.
Sources
Related posts
More in Developers
- MiniMax H3 on SGLang: 8 x B200, FL2VA server and /v1/videos
MiniMax documents a preview SGLang server for H3: 8 x B200, Ulysses 8, about 19 s per clip. The setup, the memory numbers, and when to host instead.
- MiniMax task statuses vs Sume job statuses: a mapping table
MiniMax reports queued, running, succeeded, failed, cancelled. Sume job status uses queued, processing, completed, failed, canceled. Map them correctly.
- Mirage Tesseract needs local files; Sume needs public HTTPS URLs
Mirage Tesseract runs on local files in an agent environment. Sume avatar and lip-sync inputs must be public HTTPS URLs. How to hand clips between the two.
- Music API has no duration field: steer length in the prompt
Sume's music router rejects duration and duration_seconds. Ask for a 30-second or 2-minute track in the prompt, with section timestamps, and verify.
Written by Sume