H3 on vLLM-Omni: 87 s on 4 B300 for one clip, and the cost

MiniMax says an 8.7 s H3 clip takes about 87 s on 4 B300 GPUs. Turn that into GPU-seconds, a break-even rate, and compare with a 9 s Sume job.

5 min readSume
All posts

MiniMax's H3 page gives one concrete serving figure: on the official vLLM-Omni recipe, an 8.7 second clip end to end in about 87 seconds on 4 B300 GPUs. That is 348 GPU-seconds for one clip. A nine-second hosted job on Sume at 768p costs $0.675, so self-hosting wins on raw GPU rent only if you pay less than about $6.98 per GPU-hour and keep the rig busy. All inputs below are stated, and the GPU rates are assumptions, not vendor prices.

The arithmetic

GPU-seconds per clip = 87 seconds x 4 GPUs = 348. Cost per clip = 348 / 3,600 x the hourly rate of one GPU. The table uses four hourly rates you can replace with your own quote.

Cost of one 8.7 s H3 clip at assumed GPU rates vs a 9 s Sume minimax-h3 job ($0.675), as of 2026-10-08
Assumed rate per GPU-hourRental cost for 348 GPU-secondsAgainst Sume 9 s at 768p
$2.00$0.1933Cheaper
$4.00$0.3867Cheaper
$6.00$0.58Cheaper
$8.00$0.7733More expensive

What the table leaves out

The figure is a single clip on a warm server. It does not include model load time, idle GPUs between requests, the cost of the 4-GPU node being rented in whole units, storage for the checkpoint, or the engineer who runs the stack. If the node sits idle half the time, double the cost column.

The vendor figure is also for open weights, which generate natively at 768p short side according to the same page. The 2K output is described as API only, so a 2K comparison cannot be made from local weights.

Where a hosted job is simpler

Sume's minimax-h3 accepts 5 to 15 seconds at native 480p or 768p, and its video docs describe the flow: submit to /v1/videos, get a job id and a polling URL, then download when the status is completed. There is no cluster to size. Sume reserves its price at submit, so the cost is known before the work starts.

Sume does not publish a latency promise in the pages cited here; its best-practices note says video generation usually takes from 30 seconds to several minutes depending on model and parameters. Compare that to your own measured number, not to a vendor demo.

Decision rule

Use the break-even rate as a screen, then add your idle factor.

  • If your quote is above $6.98 per GPU-hour, hosted is cheaper before idle time and labor.
  • If you pay far less than that and generate clips all day, the rig can win, but only inside the license limits: the weights exclude the US, EU, UK and South Korea.
  • If clip volume is bursty, hosted jobs cost nothing while idle.

Reading the vendor figure carefully

The figure is MiniMax's own example on its H3 deployment page, given as 'e.g.' for the official vLLM-Omni recipe, which serves an OpenAI-compatible video endpoint. SGLang Diffusion is the second serving stack the page names. The page does not say what resolution or settings produced 87 seconds, so treat 348 GPU-seconds as an order of magnitude for planning, then measure your own prompt, length and resolution before you commit to a rig.

If your clips are longer than 8.7 seconds, do not scale the number linearly without testing. Cost per output second is the figure to track: 348 GPU-seconds over 8.7 seconds is 40 GPU-seconds per output second, by division.

Sources

Related posts

More in Models

All Models posts

Written by Sume