MiniMax-H3 Fun ControlNet Union 2.0: 8 conditions vs Sume references

MiniMax-H3-Fun-Controlnet-Union-2.0 adds Scribble, Layout and Gray to five older conditions. Sume's minimax-h3 ids take image, video and audio references.

4 min readSume
All posts

MiniMax-H3-Fun-Controlnet-Union-2.0 is a control branch for MiniMax H3 that handles eight conditions with one checkpoint: Canny, Depth, HED, MLSD, Pose, Scribble, Layout and Gray, plus video inpainting. Sume does not host it. Its hosted minimax-h3 ids take reference inputs instead, and none of them is a depth, pose or edge control input.

The checkpoint facts are from its Hugging Face README and the ComfyUI changelog (September 29, 2026 entry), read 2026-10-01. Sume's are from the Video generation docs.

What is new in version 2.0?

The README compares it with v1: five conditions (Canny, Depth, HED, MLSD, Pose) become eight with Scribble, Layout and Gray; control blocks go from 5 to 10; and the inpainting recipe changes to post_norm. It warns that loading the v1 config against this checkpoint is a silent failure and says to use minimax_h3_control_inpaint_post_norm.yaml.

v1 and 2.0 as the README describes them, read 2026-10-01.
v12.0
Control conditions58 (adds Scribble, Layout, Gray)
Control blocks510
Checkpoint sizeabout 6.8 GBabout 13.5 GB
Guidanceguidance_scale = 1.0guidance_scale = 1.0

What can a hosted minimax-h3 call take instead?

Sume's catalog entries advertise supported_input_references of image_url, video_url and audio_url for the models that take them. minimax-h3-max supports text-to-video, first/last-frame image-to-video and reference-to-video. Audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max.

Is a reference the same as a control video?

Not by anything the docs say. A control video such as a depth or pose map conditions structure frame by frame in the checkpoint. Sume's docs describe references as inputs the model honors, with no mention of depth, edge, pose or layout maps, and no inpainting mask for video. For motion from a source clip, see motion transfer with a reference video.

Which route fits which job?

If you need one of the eight control conditions or video inpainting, the checkpoint runs where you run MiniMax H3 yourself, for example in ComfyUI, which lists it as supported. If a reference image, clip or audio file is enough, the hosted call needs no GPUs; compare the two in ComfyUI vs the hosted API. The checkpoint's license field is "other", pointing at the MiniMax H3 Community License.

Sources

Related posts

More in Models

All Models posts

Written by Sume