QuantFunc INT4 MiniMax H3: 3.2 s per step on an RTX 4090
QuantFunc's 4-bit MiniMax H3 claims 3.2 s per step on an RTX 4090. That is not a clip time. What the card says, what it omits, and when to use a hosted job.

The QuantFunc 4-bit build of MiniMax H3 runs at 3.2 seconds per sampling step on an RTX 4090 for a 768x768, 124-frame clip, which its card says is 3.19 times faster than an FP8 baseline at 10.2 seconds. That is a per-step number, not a per-clip time, and the card does not state a VRAM requirement, so do not size a purchase from it.
Everything about QuantFunc below is from its model card, read 2026-10-03, where it is among the models trending on Hugging Face's text-to-video list. Sume's side is from the video docs.
What does the QuantFunc build claim?
The card describes a 4-bit quantized MiniMax H3 that generates video with audio from text, images or video references. It says it runs on every NVIDIA GPU of compute capability SM75 or newer: the RTX 20, 30, 40 and 50 series, A100, H100, H200, B100, B200 and GB300. The speed claim is a measurement on one card: on an RTX 4090 at 768x768 and 124 frames, INT4 takes 3.2 seconds per step against 10.2 seconds for the FP8 baseline.
The card also reports quality against the BF16 original: about 23.7 dB PSNR, and a maximum relative output error of about 7.58e-4 from FP16 folding. PSNR of that size says outputs are close but not identical to BF16, so do not expect the same frame for the same seed.
| Item | Card value | Not stated |
|---|---|---|
| Precision | 4-bit (INT4) | Which layers stay higher precision |
| GPUs | NVIDIA SM75 and newer, RTX 20 to 50 series and data-centre cards | Minimum VRAM |
| Speed on RTX 4090 | 3.2 s per step vs 10.2 s FP8 (768x768, 124 frames) | Whole-clip time |
| Fidelity vs BF16 | About 23.7 dB PSNR; max relative error about 7.58e-4 | Perceptual quality |
| Runtime | ComfyUI-QuantFunc or the QuantFunc engine; not standard Diffusers | Other runtimes |
| Terms | MiniMax H3 Community License applies to weights and derivatives | Commercial grant beyond that licence |
How long is a clip, if a step takes 3.2 seconds?
The card gives no whole-clip time, so any figure is arithmetic on its per-step number and leaves out the text encoder, the audio path, the decode and model loading. As an illustration only: at 3.2 seconds per step, 8 steps is 25.6 seconds and 20 steps is 64 seconds of sampling at 768x768 and 124 frames on that card. Longer clips and higher resolutions cost more per step; the card does not say how much.
Measure end to end on your card and your workflow before you compare it with a hosted job. MiniMax H3's ComfyUI route against a hosted API lists the limits that change between the two.
Where does a hosted job still win?
Three places. Setup: the card says QuantFunc needs its own ComfyUI node or engine and does not work with standard Diffusers, which is one more runtime to maintain. Licence: the MiniMax H3 Community License covers the weights and derivatives, so the revenue and territory conditions still apply to your output. Scale: a card renders one clip at a time, while a hosted job queue takes many.
In Sume's Video generation docs, minimax-h3 accepts 5 to 15 seconds at native 480p or 768p, and each model lists pricing_skus in GET /v1/videos/models. Read the live price there rather than from a blog post, and compare it with your own cost per GPU hour divided by clips per hour.
Should you try the 4-bit build?
Try it if you already own an NVIDIA card in the supported range and want local iteration, and run a fixed prompt through the 4-bit build and your current precision to see the difference yourself. Skip it if you need exact reproducibility against BF16 outputs, or if you deliver a few clips a week, where a hosted job is cheaper than the hours you spend on setup.
Keep a record of the build and the settings you used, because a quantized checkpoint changes output compared with the original model and you will want to explain a difference later.
Sources
Related posts
More in Models
- Reference audio for AI video: which models accept a voice clip
Seedance 2.5, Wan 3.0 and MiniMax H3 take reference audio on Sume; Kling 3, Gemini Omni Flash and the swap rows do not. Limits and the one-reference rule.
- Omni Flash vs H3 Max for reference-to-video: limits side by side
Gemini Omni Flash 1.1 takes 10 images and 3 short videos; MiniMax H3 Max adds audio refs, 12 files total. Limits, tags and a sample request on Sume.
- Runway Veo negativePrompt: 1,000 characters, and Sume's options
Runway's API takes an optional negativePrompt of up to 1,000 characters on three Veo models. Sume's video request lists no such field, so use positive wording.
- Seedance 2.5 timestamp-level editing: what Sume exposes today
ByteDance describes timestamp-level editing and 30-second clips for Seedance 2.5. See which of those fields Sume's Video Router accepts for seedance-2.5.
Written by Sume