MiniMax H3 VRAM: what 8 GB, 24 GB and B200 setups actually are
MiniMax H3 VRAM requirements are a ladder: 8 GB with NF4 offloading, 12 to 16 GB with GGUF, 24 GB with INT8, 8 x B200 for the reference server. Which tier fits.

MiniMax H3 has no single VRAM requirement. MiniMax's integration index maps hardware to stacks, from an 8 GB minimum with NF4 and heavy offloading up to H200 and B200 cards, and its self-hosting guide uses eight B200s for the official server. The smaller tiers rely on community quantizations, and the guide says not to compare their output with lossless SGLang output.
What does each tier use?
The index pairs each hardware class with a stack. Treat each row as a listing, not a guarantee.
| Hardware | Stack named | Source |
|---|---|---|
| 8 GB minimum | DiffSynth NF4 with heavy offloading | integration index |
| 12 to 16 GB | pruned Q4_K_M GGUF, or nvfp4 with an fp8mix VAE | integration index |
| 24 GB NVIDIA | pruned INT8 ConvRot DiT, nvfp4 text encoder, ComfyUI-Easy | integration index |
| 4 x H200 | SGLang, 50 steps at 1344 x 768, 75.10 s mean | self-host guide |
| 8 x B200 | SGLang reference server, BF16/FP32, about 108 GB disk | self-host guide |
Does more memory change the clip?
On the datacenter tier, the guide reports peak VRAM of 83,578 MB per GPU for FL2VA on 8 x B300 in BF16, and says FP8 cuts about 32 GB without a significant latency change. The cards on the smaller tiers run different weights and different workflows. Keep the clip settings, the file and the loader written down next to every result.
What is the practical advice?
Choose the tier from the volume you need, not from the smallest card that technically starts. A 24 GB card with INT8 or NVFP4 files is the realistic entry point for testing. A 4 x H200 or 8 x B200 node is a production server with a production bill. The middle ground, a single 12 to 16 GB card, is for experiments with the GGUF files.
If you cannot say how many clips a week you need, use the hosted route first. Its cost is per clip, it scales to zero, and it gives you a baseline to compare a local setup against later.
When should a team call an API instead?
If the target is a 24 GB card and a handful of clips a week, the hosted route is the cheaper experiment. Sume lists minimax-h3 for 5 to 15 seconds at 480p or 768p, and minimax-h3-max for 480p, 768p and 1080p; read their price from the pricing_skus field on GET /v1/videos/models. The video docs have the request.
If you need 8 GB hardware to run at all, expect a long wait per clip; the sources I read give no timing for that tier.
Sources
Related posts
More in Models
- Mistral Large 4 on a 12-turn storyboard: tokens are 4% of the bill
Mistral Large 4 lists $1.36 input and $4.18 output per million tokens. A 12-turn run that makes six Wan 3.0 clips spends 4.3% of its total on tokens.
- Nano Banana 2.1 0.5K: Google says unsupported, Sume lists 512
Google's docs say the 512px tier is not supported on Nano Banana 2.1, yet the Sume catalog lists 512. Which to trust, and how to check the row live.
- Nano Banana 2.1 16:9 output size: 1376x768 up to 5504x3072
What pixel size a 16:9 Nano Banana 2.1 image comes out at per tier, and the 15 px crop that turns the 2K frame into exact 1920x1080.
- Nano Banana 2.1 1K image is 1,120 tokens: Google's $0.0336 explained
Google prices Nano Banana 2.1 image output at $30 per million tokens: a 1K image is 1,120 tokens, $0.0336. 2K and 4K implied token counts, Batch half price.
Written by Sume