QuantFunc INT4 MiniMax H3: 3.2 s per step on an RTX 4090

QuantFunc's 4-bit MiniMax H3 claims 3.2 s per step on an RTX 4090. That is not a clip time. What the card says, what it omits, and when to use a hosted job.

5 min readSume
All posts

The QuantFunc 4-bit build of MiniMax H3 runs at 3.2 seconds per sampling step on an RTX 4090 for a 768x768, 124-frame clip, which its card says is 3.19 times faster than an FP8 baseline at 10.2 seconds. That is a per-step number, not a per-clip time, and the card does not state a VRAM requirement, so do not size a purchase from it.

Everything about QuantFunc below is from its model card, read 2026-10-03, where it is among the models trending on Hugging Face's text-to-video list. Sume's side is from the video docs.

What does the QuantFunc build claim?

The card describes a 4-bit quantized MiniMax H3 that generates video with audio from text, images or video references. It says it runs on every NVIDIA GPU of compute capability SM75 or newer: the RTX 20, 30, 40 and 50 series, A100, H100, H200, B100, B200 and GB300. The speed claim is a measurement on one card: on an RTX 4090 at 768x768 and 124 frames, INT4 takes 3.2 seconds per step against 10.2 seconds for the FP8 baseline.

The card also reports quality against the BF16 original: about 23.7 dB PSNR, and a maximum relative output error of about 7.58e-4 from FP16 folding. PSNR of that size says outputs are close but not identical to BF16, so do not expect the same frame for the same seed.

QuantFunc card facts, read 2026-10-03.
ItemCard valueNot stated
Precision4-bit (INT4)Which layers stay higher precision
GPUsNVIDIA SM75 and newer, RTX 20 to 50 series and data-centre cardsMinimum VRAM
Speed on RTX 40903.2 s per step vs 10.2 s FP8 (768x768, 124 frames)Whole-clip time
Fidelity vs BF16About 23.7 dB PSNR; max relative error about 7.58e-4Perceptual quality
RuntimeComfyUI-QuantFunc or the QuantFunc engine; not standard DiffusersOther runtimes
TermsMiniMax H3 Community License applies to weights and derivativesCommercial grant beyond that licence

How long is a clip, if a step takes 3.2 seconds?

The card gives no whole-clip time, so any figure is arithmetic on its per-step number and leaves out the text encoder, the audio path, the decode and model loading. As an illustration only: at 3.2 seconds per step, 8 steps is 25.6 seconds and 20 steps is 64 seconds of sampling at 768x768 and 124 frames on that card. Longer clips and higher resolutions cost more per step; the card does not say how much.

Measure end to end on your card and your workflow before you compare it with a hosted job. MiniMax H3's ComfyUI route against a hosted API lists the limits that change between the two.

Where does a hosted job still win?

Three places. Setup: the card says QuantFunc needs its own ComfyUI node or engine and does not work with standard Diffusers, which is one more runtime to maintain. Licence: the MiniMax H3 Community License covers the weights and derivatives, so the revenue and territory conditions still apply to your output. Scale: a card renders one clip at a time, while a hosted job queue takes many.

In Sume's Video generation docs, minimax-h3 accepts 5 to 15 seconds at native 480p or 768p, and each model lists pricing_skus in GET /v1/videos/models. Read the live price there rather than from a blog post, and compare it with your own cost per GPU hour divided by clips per hour.

Should you try the 4-bit build?

Try it if you already own an NVIDIA card in the supported range and want local iteration, and run a fixed prompt through the 4-bit build and your current precision to see the difference yourself. Skip it if you need exact reproducibility against BF16 outputs, or if you deliver a few clips a week, where a hosted job is cheaper than the hours you spend on setup.

Keep a record of the build and the settings you used, because a quantized checkpoint changes output compared with the original model and you will want to explain a difference later.

Sources

Related posts

More in Models

All Models posts

Written by Sume