MiniMax H3 in ComfyUI or through a hosted API: what differs

ComfyUI documents three MiniMax H3 workflows with local model files; a hosted API returns a job. The limits that change between the two, from the docs.

4 min readSume
All posts

ComfyUI runs MiniMax H3 from local model files in text-to-video, image-to-video and reference-to-video workflows, while a hosted API such as Sume's POST /v1/videos takes the same three modes as one JSON request and returns a job. Pick ComfyUI when you need graph-level control, and the API when you need to call it from code without GPUs.

ComfyUI facts are from its H3 tutorial and MiniMax's model card; Sume facts are from the Video generation docs, all read 2026-09-29.

What does ComfyUI need for H3?

Per its tutorial: a diffusion model variant (fl2va for text and image-to-video, ref2va for reference-to-video), a text encoder, a video VAE and an audio VAE. Optional turbo LoRAs cut sampling to 8 steps for text and image-to-video and 4 steps for reference-to-video, against a 20-step default. The tutorial also documents a MiniMaxH3AddGuide node that anchors images or audio at any frame, and latent noise masks for regenerating part of a clip.

How do the limits compare?

Limits as stated by each source, read 2026-09-29.
TopicComfyUI and model cardSume `minimax-h3`
Native canvas768 px short edge, sizes rounded to a multiple of 32Native 480p and 768p; 2K and 4K are priced upscales
LengthDuration snaps to a 17-frame-per-block grid at 24 fps5–15 seconds, whole seconds
ReferencesUp to 9 images, 3 videos, 3 audio clipsImage, video and audio types; counts are not in the docs
Where it runsYour hardwareSume's provider, billed per output second

Can I mask and regenerate part of a clip through the API?

Sume does not list that. Its H3 ids take a prompt, frames or references; the latent-mask editing in ComfyUI is a graph feature. If you need to change one region of a clip in Sume, the docs list gemini-omni-flash-1.1's video_url edit mode through the Video Router, a different model.

Which one should I start with?

If you have not run H3 before, one hosted request shows what the model does with your prompt before you set up drivers, encoders and VAEs. If you already have the GPUs and want turbo LoRAs or guide anchoring, start in ComfyUI. Check the license on the model card either way.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume