MiniMax-H3 Fun ControlNet Union 2.0: 8 conditions vs Sume references
MiniMax-H3-Fun-Controlnet-Union-2.0 adds Scribble, Layout and Gray to five older conditions. Sume's minimax-h3 ids take image, video and audio references.

MiniMax-H3-Fun-Controlnet-Union-2.0 is a control branch for MiniMax H3 that handles eight conditions with one checkpoint: Canny, Depth, HED, MLSD, Pose, Scribble, Layout and Gray, plus video inpainting. Sume does not host it. Its hosted minimax-h3 ids take reference inputs instead, and none of them is a depth, pose or edge control input.
The checkpoint facts are from its Hugging Face README and the ComfyUI changelog (September 29, 2026 entry), read 2026-10-01. Sume's are from the Video generation docs.
What is new in version 2.0?
The README compares it with v1: five conditions (Canny, Depth, HED, MLSD, Pose) become eight with Scribble, Layout and Gray; control blocks go from 5 to 10; and the inpainting recipe changes to post_norm. It warns that loading the v1 config against this checkpoint is a silent failure and says to use minimax_h3_control_inpaint_post_norm.yaml.
| v1 | 2.0 | |
|---|---|---|
| Control conditions | 5 | 8 (adds Scribble, Layout, Gray) |
| Control blocks | 5 | 10 |
| Checkpoint size | about 6.8 GB | about 13.5 GB |
| Guidance | guidance_scale = 1.0 | guidance_scale = 1.0 |
What can a hosted minimax-h3 call take instead?
Sume's catalog entries advertise supported_input_references of image_url, video_url and audio_url for the models that take them. minimax-h3-max supports text-to-video, first/last-frame image-to-video and reference-to-video. Audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max.
Is a reference the same as a control video?
Not by anything the docs say. A control video such as a depth or pose map conditions structure frame by frame in the checkpoint. Sume's docs describe references as inputs the model honors, with no mention of depth, edge, pose or layout maps, and no inpainting mask for video. For motion from a source clip, see motion transfer with a reference video.
Which route fits which job?
If you need one of the eight control conditions or video inpainting, the checkpoint runs where you run MiniMax H3 yourself, for example in ComfyUI, which lists it as supported. If a reference image, clip or audio file is enough, the hosted call needs no GPUs; compare the two in ComfyUI vs the hosted API. The checkpoint's license field is "other", pointing at the MiniMax H3 Community License.
Sources
Related posts
More in Models
- MiniMax H3 license for an EU company: contact MiniMax first
MiniMax H3 license Section II invites people in the EU, UK, Korea and USA to contact MiniMax about a license. What it promises and omits.
- MiniMax H3 license: EU, UK, Korea and US are Excluded Territories
The MiniMax H3 Community License defines the EU, UK, Republic of Korea and USA as Excluded Territories and says use there is not authorized. What it quotes.
- minimax/hailuo-3 on OpenRouter vs Sume minimax-h3 id
OpenRouter names the model minimax/hailuo-3. Sume uses bare catalog ids, so the matching id is minimax-h3, native 480p/768p, with 2K and 4K as priced upscales.
- Muse Voice Transcribe 25+ languages and code-switching vs Sume STT
Muse Voice Transcribe supports 25+ languages and code-switching in one conversation. For mixed-language audio on Sume STT, omit language_code or set one hint.
Written by Sume