MiniMax H3 on SGLang: 8 x B200, FL2VA server and /v1/videos

MiniMax documents a preview SGLang server for H3: 8 x B200, Ulysses 8, about 19 s per clip. The setup, the memory numbers, and when to host instead.

5 min readSume
All posts

MiniMax's self-hosting guide documents a preview SGLang server for H3 on one node of 8 NVIDIA B200 GPUs, with Ulysses parallelism of 8 and the official mixed BF16/FP32 weights. Its reference timing is about 19 seconds for a 1344 x 768, 124-frame text-to-video-with-audio run at 50 steps, with peak memory around 84 GB per GPU.

Everything below is from the Run and Self-host MiniMax H3 guide, read on 2026-10-02. It calls the SGLang route a preview needing SGLang v0.5.17 or newer, so treat the numbers as the vendor's reference, not a guarantee for your hardware. Hosted Sume access to the same family is described in Video generation.

What does the vendor setup look like?

The guide pins an SGLang commit, creates a Python 3.12 environment with uv, installs the diffusion extra and downloads only the model_index.json and FL2VA/* files of the MiniMaxAI/MiniMax-H3 repo at a pinned revision. SGLang selects the checkpoint with --model-variant fl2va.

The server binds to 127.0.0.1 on port 30010 in the example, and readiness is a call to /health that should answer {"status":"ok"}. A video request is a POST /v1/videos with task, target (short edge, aspect ratio, duration) and sampling fields; you then poll and fetch /v1/videos/{id}/content.

sglang serve \
  --model-path "$MINIMAX_H3_MODEL_DIR" \
  --model-variant fl2va \
  --num-gpus 8 --tp-size 1 \
  --ulysses-degree 8 \
  --encoder-parallel auto \
  --performance-mode speed \
  --host 127.0.0.1 --port 30010

What does it cost in memory and time?

The page gives one reproducible baseline and one lighter option. It does not give a price per clip, so cost per video depends on what you pay for eight GPUs.

SGLang reference numbers from MiniMax, read 2026-10-02
ItemValue
Hardware8 x NVIDIA B200, one node
ParallelismUlysses degree 8
Text-to-video latencyAbout 19 s at 1344 x 768, 124 frames, 50 steps
Peak VRAM, BF16About 84 GB per GPU
Peak VRAM, FP8About 51 GB per GPU, approximate quantization
StatusPreview, SGLang v0.5.17 or newer

What are the license and security limits?

The guide states that H3 is governed by the MiniMax H3 Community License Agreement, and that the United States, European Union, United Kingdom and South Korea are Excluded Territories requiring a separate license. Read that before you deploy; the Sume posts on the territory rule cover what it means for outputs.

For exposure beyond localhost the guide asks for an authenticated reverse proxy, TLS at ingress, a firewalled backend port, request quotas, concurrency limits, audit logging, read-only reference media, and for remote URLs a domain allowlist with redirect, timeout, byte and duration limits. Those are real operating costs of running your own endpoint.

When is hosting still the easier path?

Self-hosting makes sense when you need control of weights, a LoRA, or data staying inside your network. If you only need clips, a hosted API removes the cluster, the pinned commits and the license review of your own deployment.

On Sume, minimax-h3 accepts 5 to 15 seconds at native 480p or 768p, minimax-h3-max adds 1080p as a latent refinement, and billing is provider list times 1.25 on the workspace balance. Submit to POST /v1/videos, poll the job, and keep the job id.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume