MiniMax H3 on a 12 GB RTX 3060: 480p only, 42 GB, and the price

MiniMax says 12 GB of VRAM runs 480p H3 clips with audio from a pruned int8 checkpoint of about 42 GB. What that buys, and a 480p clip on Sume costs.

5 min readSume
All posts

MiniMax's own H3 FAQ says a 12 GB card of RTX 3060 class can run 480p clips with audio using the pruned int8 checkpoint, which is about 42 GB of downloads. Larger cards are more comfortable: the page puts 16 to 24 GB at 'minutes per clip once optimized'. For comparison, a five-second 480p clip on Sume's minimax-h3 costs $0.3125 (5 seconds x $0.05 list x 1.25), and the license behind the local weights excludes the US, EU, UK and South Korea. Figures are from the vendor page read on 2026-10-08.

What MiniMax lists by hardware

The FAQ calls these 'community-verified reference points', not guarantees from MiniMax. Treat them as the starting point for your own test, and treat the cost of the first failed run as part of the project.

H3 open-weights hardware reference points, MiniMax H3 FAQ read 2026-10-08
HardwareWhat the page says
12 GB VRAM, RTX 3060 class480p clips with audio, pruned int8 checkpoint, about 42 GB of downloads
16 to 24 GB VRAMComfortable; minutes per clip once optimized
RTX 5090 or DGX SparkRoughly 4x further with NVIDIA's Sol Engine stack
Apple SiliconMLX with 128 GB unified memory recommended, or an NF4 route from 16 GB
System RAMPlan 32 to 64 GB; layers that do not fit in VRAM are staged from RAM

What 480p means for the output

The open weights generate natively at 768p on the short side, the page says, and the 2K output relies on API-only stages. A 12 GB setup produces 480p, the lowest tier. If your end product is a vertical ad that has to look sharp on a phone, a 480p draft is a draft. That is fine for prompt testing, which is the use case where a local card makes most sense.

The same page lists the open variants as FL2VA (text-to-video with optional first and last frame) and Ref2VA (omni-reference with images, videos and audio). Each is a separate checkpoint, so the 42 GB figure covers one route, not both.

Cost of 480p drafts on Sume

Sume prices minimax-h3 per output second at 480p. Sume's docs say the model accepts 5 to 15 seconds at native 480p or 768p. The table shows what 20 draft clips cost at different lengths.

Sume minimax-h3 at 480p, $0.0625 per second (list $0.05 x 1.25), as of 2026-10-08
Clip lengthPer clip20 draft clips
5 s$0.3125$6.25
8 s$0.50$10.00
10 s$0.625$12.50
15 s$0.9375$18.75

When the local card still wins

A card you already own has no per-clip cost, so a few hundred 480p tests a month could be cheaper locally. Add the 42 GB download, the setup time and the power cost, and the break-even moves toward hosted for most small teams outside the excluded regions.

  • Choose local if you already own the GPU, are outside the excluded regions, and want to iterate on prompts at 480p.
  • Choose hosted if you need 768p or higher, need the same result tomorrow from a laptop, or are in the US, EU, UK or South Korea.
  • Sume's job flow is submit, poll or webhook, then download, so a script that works for one clip works for twenty.

Other requirements on the same page

The same FAQ says the fastest path for individuals is ComfyUI, with day-0 support and bundled workflow templates, and names the Diffusers release and DiffSynth-Studio for Python pipelines. It lists WanGP for low-VRAM machines and says community stacks go as low as 5 to 8 GB. Those are vendor-cited community routes, so the quality at 5 GB is not something this post can vouch for.

Sources

Related posts

More in Models

All Models posts

Written by Sume