MiniMax H3 on 4 H200s: 75 s, or 53.7 s with Cache-DiT, and break-even

MiniMax's guide times one H3 run on 4 H200s at 75.1 s, about 53.7 s with Cache-DiT. The GPU-seconds, and the hourly rate where renting breaks even.

5 min readSume
All posts

MiniMax's self-hosting guide gives one 4 x NVIDIA H200 data point: a text-to-video run at 50 steps and 1344 x 768 averages 75.10 seconds, and about 53.70 seconds with Cache-DiT in quality mode. That is 300.4 and about 214.8 GPU-seconds. A hosted 5-second 768p minimax-h3 clip on Sume costs $0.375, so renting only wins below roughly $4.49 per GPU-hour at full utilization, and below about $6.28 with Cache-DiT.

How is that break-even computed?

Sume reserves provider list times 1.25. The code lists minimax-h3 at $0.06 per second at 768p, so the Sume price is $0.075 per second, or $0.375 for 5 seconds. The guide does not state the clip length for its timing, so I assume 5 seconds; if yours was longer, the real break-even is higher. The arithmetic: GPU-seconds divided by 3,600 is GPU-hours per clip, and $0.375 divided by that is the break-even rate.

Break-even GPU rate for one 5-second 768p clip (arithmetic on the guide's timings, read 2026-10-09)
RunSeconds on 4 GPUsGPU-secondsGPU-hours per clipBreak-even $/GPU-hour
4 x H200, 50 steps75.10300.40.08344.49
4 x H200, Cache-DiT quality modeabout 53.70about 214.80.05976.28

What does the break-even leave out?

It assumes the GPUs are busy every second you pay for. The guide's number is one run with the model already resident, so it excludes load time, idle time, failed runs and your engineering hours. It is also a text-to-video run in BF16; the guide separately says Cache-DiT times are approximate.

If your clips are 15 seconds, check whether the same timing still holds before you scale the numbers up. I did not find a duration-versus-time table in the guide.

  • Idle GPU-hours push the real break-even down fast.
  • A reservation or spot interruption changes the rate, not the formula.
  • The license excludes the US, EU, UK and South Korea.

Worked example with a made-up rate

Suppose your provider charges $3.00 per GPU-hour. One 75.10-second run on 4 GPUs is 0.0834 GPU-hours, which is about $0.25 at full utilization, below the $0.375 hosted price for a 5-second 768p clip. At 50 percent utilization the same clip costs $0.50, and the hosted price wins. The $3.00 is only an example; use your own quote.

That is why utilization matters more than the benchmark. A box that runs clips all day can beat a per-second price; a box that waits for requests cannot. If your load comes in bursts, the hosted job is paid only while it runs.

What is the rule of thumb?

Ask for your price per GPU-hour, multiply by 0.0834, and compare with $0.375. Then multiply your expected idle share on top. Read live prices from GET /v1/videos/models rather than from this post; see the video docs.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume