MiniMax H3 vs H3 Max on Sume: 768p costs 33 percent more on Max

At 768p a second of minimax-h3-max costs $0.10 on Sume against $0.075 for minimax-h3. What the extra third buys, with a 10-second price table.

4 min readSume
All posts

The gap

Both models are MiniMax H3 family entries in the Sume catalog, both take 5 to 15 seconds, and at 480p they cost the same: $0.0625 a second. At 768p they split. minimax-h3 is $0.075 a second and minimax-h3-max is $0.10, which is 33 percent more. Sume applies the same 1.25 factor to both, so the gap is in the provider list ($0.06 against $0.08 per second in Sume's pricing table).

What the extra money buys

The Video Router docs describe minimax-h3-max as the faster 768p variant. It does text-to-video, start and end frame image-to-video, and reference-to-video through reference_*_urls, at 480p, 768p and 1080p for 5 to 15 seconds, with native stereo audio. Its 1080p is a latent refinement from native 768p, not native 1080p. minimax-h3 is native 480p and 768p, with 2K and 4K upscales billed if requested.

So the choice is not only price: H3 Max has the 1080p row and the reference-to-video route; H3 is cheaper at 768p and reaches 2K or 4K through upscaling.

10-second job on Sume by tier (list x 1.25; read 2026-10-05)
Model480p768pHigher tier
minimax-h3 (10 s)$0.625$0.752K $1.625, 4K $2.00
minimax-h3-max (10 s)$0.625$1.001080p $2.00

A simple rule

  • Drafting at 480p: either model, same price, so pick the one you will finish with.
  • Final at 768p with no references: minimax-h3 saves $0.25 on a 10-second clip.
  • Final that needs references or a 1080p output: minimax-h3-max.
  • Neither accepts seed, so a draft is a different take from the final.

Verify

Read each model's capabilities from GET /v1/video-router/models. The envelope differs per model, and the live catalog is the source of truth for resolutions and durations.

A worked comparison for a campaign

A campaign needs twenty 8-second clips at 768p. On minimax-h3 that is 20 x 8 x $0.075 = $12.00. On minimax-h3-max it is 20 x 8 x $0.10 = $16.00. The $4.00 difference is the price of the Max variant for that batch. If a third of the clips need reference-to-video, which H3 Max handles through reference_*_urls, those clips could not run on minimax-h3 at all, and the fair comparison is a mixed batch.

If the batch is drafts at 480p, the price is identical, $0.0625 a second, so the decision should be made on the final tier.

Common mistakes

Treating the Max 1080p row as native 1080p is the most common one. The docs describe it as a latent refinement from native 768p, so it is a different pipeline from a model that renders natively at 1080p. Another is assuming that either model accepts a 4-second clip. Both start at 5 seconds, so a 4-second request fails validation before any reserve is taken.

Finally, remember that each model has its own capability envelope. Read capabilities from the Video Router model endpoint rather than copying parameters from one model's example to the other.

One more way to decide is to count the retakes. If Max gets an acceptable result in one attempt where the base model needs two, the base model's cheaper rate is cancelled out: two 8-second attempts at $0.075 a second are $1.20, one attempt at $0.10 a second is $0.80. The reverse also holds, so do not assume the pricier variant is the efficient one. Run a small A/B on your own prompts, five each, read the captured amounts from the ledger, and pick on measured cost per accepted clip rather than on the rate alone.

The ledger makes this easy because every job has its own captured row. Failed jobs are refunded, so the retake count you pay for is only the successful generations.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume