MiniMax H3 GGUF sizes: Q4_K_M 18.5 GiB, IQ1_S 3.78 GiB
MiniMax's H3 integration index lists GGUF files from 3.78 to 33.56 GiB against a 61.73 GiB BF16 original. Sizes in a table, plus when to call a job instead.

MiniMax H3 has GGUF files, but they are not the main release. MiniMax's own integration index lists GGUF quantizations from IQ1_S at 3.78 GiB up to Q8_0 at 33.56 GiB, next to a 61.73 GiB BF16 original. The self-hosting guide describes the official weights as mixed BF16/FP32, so a GGUF is a smaller approximation of them, not the model you would get from the SGLang route.
Which GGUF sizes are listed?
Sizes are GiB (1024 cubed bytes), so a 18.50 GiB file is about 19.9 GB in the units a disk vendor uses.
| Format | Size on disk | What the index says |
|---|---|---|
| BF16 safetensors | 61.73 GiB | the original |
| INT8 ConvRot (pruned) | 19.53 GiB | entry point for a 24 GB card |
| FP8 Scaled | 19.52 GiB | safetensors variant |
| NVFP4 (pruned) | 10.86 to 11.67 GiB | safetensors variant |
| GGUF Q8_0 | 33.56 GiB | largest GGUF |
| GGUF Q4_K_M | 18.50 GiB | paired with 12 to 16 GB cards in its hardware table |
| GGUF Q2_K | 6.26 to 17.42 GiB | range given by the index |
| GGUF IQ1_S | 3.78 GiB | smallest |
Does a smaller file give the same video?
The index does not say. The self-hosting guide does give a rule for anyone comparing outputs: do not mix ComfyUI (quantized) and SGLang (lossless) outputs in benchmarks. Treat every quantized result as its own model when you compare cost or quality, and note which file made each clip.
How do you pick a size for your card?
Start from the hardware row in the same index. It pairs 12 to 16 GB cards with a pruned Q4_K_M GGUF or an nvfp4 file, and 24 GB NVIDIA cards with the pruned INT8 ConvRot transformer. Then leave headroom: the file size is only the weights, and the text encoder, the VAE and the activations need memory too. The index does not give a total for any of these stacks, so the first run on your card is the real measurement.
Keep a log of file name, loader and settings next to each clip. Without it, a result from Q4_K_M and a result from NVFP4 look like the same model.
When is a hosted job simpler?
A GGUF saves disk and memory. It does not remove the work of picking a file, a loader and a workflow, and it does not change the license. The hosted route on Sume is one id: minimax-h3 takes 5 to 15 seconds at 480p or 768p, with 768p as the default, through POST /v1/videos. The model's own card says 4 to 15 seconds, so a 4-second clip is something only a local run can make.
Check the current ids and prices with GET /v1/videos/models before you plan a budget; the pricing_skus field is the source of truth. The video docs show the full request.
Sources
Related posts
More in Models
- H3 Max 4K request fails: H3 takes 2K and 4K as 768p upscales
minimax-h3-max accepts 480p, 768p or 1080p only. A 4K request belongs on minimax-h3, where it is a 768p upscale at $0.20 a second, $1.60 for 8 s on Sume.
- MiniMax H3 on a Mac: the h3.c Metal port, or a hosted job
MiniMax's integration index lists h3.c, a Metal-native MIT-licensed H3 port for Apple Silicon. What the index does and does not claim, and the hosted route.
- MiniMax H3 prompting guide: the h3-prompt-writing skill
MiniMax ships nine H3 prompting skills in its repo; one is portable via npx skills add. What is public, what is Hub-only, and how to use the output on Sume.
- MiniMax H3 VRAM: what 8 GB, 24 GB and B200 setups actually are
MiniMax H3 VRAM requirements are a ladder: 8 GB with NF4 offloading, 12 to 16 GB with GGUF, 24 GB with INT8, 8 x B200 for the reference server. Which tier fits.
Written by Sume