MiniMax H3 GGUF sizes: Q4_K_M 18.5 GiB, IQ1_S 3.78 GiB

MiniMax's H3 integration index lists GGUF files from 3.78 to 33.56 GiB against a 61.73 GiB BF16 original. Sizes in a table, plus when to call a job instead.

4 min readSume
All posts

MiniMax H3 has GGUF files, but they are not the main release. MiniMax's own integration index lists GGUF quantizations from IQ1_S at 3.78 GiB up to Q8_0 at 33.56 GiB, next to a 61.73 GiB BF16 original. The self-hosting guide describes the official weights as mixed BF16/FP32, so a GGUF is a smaller approximation of them, not the model you would get from the SGLang route.

Which GGUF sizes are listed?

Sizes are GiB (1024 cubed bytes), so a 18.50 GiB file is about 19.9 GB in the units a disk vendor uses.

H3 file sizes listed in MiniMax's integration index (read 2026-10-09)
FormatSize on diskWhat the index says
BF16 safetensors61.73 GiBthe original
INT8 ConvRot (pruned)19.53 GiBentry point for a 24 GB card
FP8 Scaled19.52 GiBsafetensors variant
NVFP4 (pruned)10.86 to 11.67 GiBsafetensors variant
GGUF Q8_033.56 GiBlargest GGUF
GGUF Q4_K_M18.50 GiBpaired with 12 to 16 GB cards in its hardware table
GGUF Q2_K6.26 to 17.42 GiBrange given by the index
GGUF IQ1_S3.78 GiBsmallest

Does a smaller file give the same video?

The index does not say. The self-hosting guide does give a rule for anyone comparing outputs: do not mix ComfyUI (quantized) and SGLang (lossless) outputs in benchmarks. Treat every quantized result as its own model when you compare cost or quality, and note which file made each clip.

How do you pick a size for your card?

Start from the hardware row in the same index. It pairs 12 to 16 GB cards with a pruned Q4_K_M GGUF or an nvfp4 file, and 24 GB NVIDIA cards with the pruned INT8 ConvRot transformer. Then leave headroom: the file size is only the weights, and the text encoder, the VAE and the activations need memory too. The index does not give a total for any of these stacks, so the first run on your card is the real measurement.

Keep a log of file name, loader and settings next to each clip. Without it, a result from Q4_K_M and a result from NVFP4 look like the same model.

When is a hosted job simpler?

A GGUF saves disk and memory. It does not remove the work of picking a file, a loader and a workflow, and it does not change the license. The hosted route on Sume is one id: minimax-h3 takes 5 to 15 seconds at 480p or 768p, with 768p as the default, through POST /v1/videos. The model's own card says 4 to 15 seconds, so a 4-second clip is something only a local run can make.

Check the current ids and prices with GET /v1/videos/models before you plan a budget; the pricing_skus field is the source of truth. The video docs show the full request.

Sources

Related posts

More in Models

All Models posts

Written by Sume