FastH3 on vLLM-Omni: a 10-second H3 video in 8.7 seconds

vLLM's team reports 10.1 seconds of H3 video and audio in about 8.7 s on 8 B300 GPUs. What the number covers, and what Sume's hosted minimax-h3 does.

5 min readSume
All posts

The vLLM-Omni team reports that MiniMax H3 with FastVideo's FastH3 adapter produced a 10.125 second video-and-audio file in 8.678 to 8.710 seconds on eight NVIDIA B300 GPUs (read 2026-10-02). That is a self-reported team benchmark on one hardware setup, not a number you will see from Sume's hosted minimax-h3, which runs as an asynchronous job that typically takes from 30 seconds to several minutes.

What exactly was measured?

The post defines real time narrowly: the whole response is ready faster than its playback length. It does not mean streaming or time to first frame. FastH3 is a four-step student of H3, replacing 49 transformer forwards with four, while reusing the original encoder, VAEs and tokenizers.

FastH3 timings on 8x B300, 1344x768 at 24 FPS (read 2026-10-02)
Clip lengthEnd-to-end time reported
5 seconds4.396 to 4.602 s
10 seconds8.678 to 8.710 s
15 seconds14.059 to 14.177 s

What are the caveats?

  • FastH3 v1 handles text-to-video-audio only, not first/last-frame or reference tasks.
  • It cannot be combined with some offload, quantization and attention options.
  • Quality parity against base H3 across several seeds is listed as pending.
  • The model runs under the MiniMax H3 Community License Agreement, which the post says needs legal review for commercial use, hosted services and territory.

What does this change for Sume users?

Nothing in the catalog. Sume's video router docs list minimax-h3 for 5 to 15 seconds with text-to-video, first and last frame, and reference inputs, and the docs describe no distilled four-step variant. You cannot pick FastH3 on Sume, and Sume does not promise faster-than-playback rendering.

Where it matters is capacity planning. If you were weighing self-hosting open H3 weights for latency, the post shows that interactive speed needs a multi-GPU Blackwell box and gives up reference inputs in v1. For most teams the hosted route, billed per second at catalog list times 1.25, stays simpler.

How to read the claim

Wait for an independent reproduction before planning around 8.7 seconds, and note the resolution: 1344 by 768 is near MiniMax's 768P tier, not 2K. If you need an interactive product, test first-frame latency separately, because the post says time to first frame was not what it measured.

Sources

Related posts

More in Models

All Models posts

Written by Sume