MiniMax H3 VRAM: what 8 GB, 24 GB and B200 setups actually are

MiniMax H3 VRAM requirements are a ladder: 8 GB with NF4 offloading, 12 to 16 GB with GGUF, 24 GB with INT8, 8 x B200 for the reference server. Which tier fits.

4 min readSume
All posts

MiniMax H3 has no single VRAM requirement. MiniMax's integration index maps hardware to stacks, from an 8 GB minimum with NF4 and heavy offloading up to H200 and B200 cards, and its self-hosting guide uses eight B200s for the official server. The smaller tiers rely on community quantizations, and the guide says not to compare their output with lossless SGLang output.

What does each tier use?

The index pairs each hardware class with a stack. Treat each row as a listing, not a guarantee.

Hardware tiers named for H3 (read 2026-10-09)
HardwareStack namedSource
8 GB minimumDiffSynth NF4 with heavy offloadingintegration index
12 to 16 GBpruned Q4_K_M GGUF, or nvfp4 with an fp8mix VAEintegration index
24 GB NVIDIApruned INT8 ConvRot DiT, nvfp4 text encoder, ComfyUI-Easyintegration index
4 x H200SGLang, 50 steps at 1344 x 768, 75.10 s meanself-host guide
8 x B200SGLang reference server, BF16/FP32, about 108 GB diskself-host guide

Does more memory change the clip?

On the datacenter tier, the guide reports peak VRAM of 83,578 MB per GPU for FL2VA on 8 x B300 in BF16, and says FP8 cuts about 32 GB without a significant latency change. The cards on the smaller tiers run different weights and different workflows. Keep the clip settings, the file and the loader written down next to every result.

What is the practical advice?

Choose the tier from the volume you need, not from the smallest card that technically starts. A 24 GB card with INT8 or NVFP4 files is the realistic entry point for testing. A 4 x H200 or 8 x B200 node is a production server with a production bill. The middle ground, a single 12 to 16 GB card, is for experiments with the GGUF files.

If you cannot say how many clips a week you need, use the hosted route first. Its cost is per clip, it scales to zero, and it gives you a baseline to compare a local setup against later.

When should a team call an API instead?

If the target is a 24 GB card and a handful of clips a week, the hosted route is the cheaper experiment. Sume lists minimax-h3 for 5 to 15 seconds at 480p or 768p, and minimax-h3-max for 480p, 768p and 1080p; read their price from the pricing_skus field on GET /v1/videos/models. The video docs have the request.

If you need 8 GB hardware to run at all, expect a long wait per clip; the sources I read give no timing for that tier.

Sources

Related posts

More in Models

All Models posts

Written by Sume