MiniMax H3 on SGLang: 8 x B200, FL2VA server and /v1/videos
MiniMax documents a preview SGLang server for H3: 8 x B200, Ulysses 8, about 19 s per clip. The setup, the memory numbers, and when to host instead.

MiniMax's self-hosting guide documents a preview SGLang server for H3 on one node of 8 NVIDIA B200 GPUs, with Ulysses parallelism of 8 and the official mixed BF16/FP32 weights. Its reference timing is about 19 seconds for a 1344 x 768, 124-frame text-to-video-with-audio run at 50 steps, with peak memory around 84 GB per GPU.
Everything below is from the Run and Self-host MiniMax H3 guide, read on 2026-10-02. It calls the SGLang route a preview needing SGLang v0.5.17 or newer, so treat the numbers as the vendor's reference, not a guarantee for your hardware. Hosted Sume access to the same family is described in Video generation.
What does the vendor setup look like?
The guide pins an SGLang commit, creates a Python 3.12 environment with uv, installs the diffusion extra and downloads only the model_index.json and FL2VA/* files of the MiniMaxAI/MiniMax-H3 repo at a pinned revision. SGLang selects the checkpoint with --model-variant fl2va.
The server binds to 127.0.0.1 on port 30010 in the example, and readiness is a call to /health that should answer {"status":"ok"}. A video request is a POST /v1/videos with task, target (short edge, aspect ratio, duration) and sampling fields; you then poll and fetch /v1/videos/{id}/content.
sglang serve \
--model-path "$MINIMAX_H3_MODEL_DIR" \
--model-variant fl2va \
--num-gpus 8 --tp-size 1 \
--ulysses-degree 8 \
--encoder-parallel auto \
--performance-mode speed \
--host 127.0.0.1 --port 30010What does it cost in memory and time?
The page gives one reproducible baseline and one lighter option. It does not give a price per clip, so cost per video depends on what you pay for eight GPUs.
| Item | Value |
|---|---|
| Hardware | 8 x NVIDIA B200, one node |
| Parallelism | Ulysses degree 8 |
| Text-to-video latency | About 19 s at 1344 x 768, 124 frames, 50 steps |
| Peak VRAM, BF16 | About 84 GB per GPU |
| Peak VRAM, FP8 | About 51 GB per GPU, approximate quantization |
| Status | Preview, SGLang v0.5.17 or newer |
What are the license and security limits?
The guide states that H3 is governed by the MiniMax H3 Community License Agreement, and that the United States, European Union, United Kingdom and South Korea are Excluded Territories requiring a separate license. Read that before you deploy; the Sume posts on the territory rule cover what it means for outputs.
For exposure beyond localhost the guide asks for an authenticated reverse proxy, TLS at ingress, a firewalled backend port, request quotas, concurrency limits, audit logging, read-only reference media, and for remote URLs a domain allowlist with redirect, timeout, byte and duration limits. Those are real operating costs of running your own endpoint.
When is hosting still the easier path?
Self-hosting makes sense when you need control of weights, a LoRA, or data staying inside your network. If you only need clips, a hosted API removes the cluster, the pinned commits and the license review of your own deployment.
On Sume, minimax-h3 accepts 5 to 15 seconds at native 480p or 768p, minimax-h3-max adds 1080p as a latent refinement, and billing is provider list times 1.25 on the workspace balance. Submit to POST /v1/videos, poll the job, and keep the job id.
Sources
Related posts
More in Developers
- MiniMax task statuses vs Sume job statuses: a mapping table
MiniMax reports queued, running, succeeded, failed, cancelled. Sume job status uses queued, processing, completed, failed, canceled. Map them correctly.
- Mirage Tesseract needs local files; Sume needs public HTTPS URLs
Mirage Tesseract runs on local files in an agent environment. Sume avatar and lip-sync inputs must be public HTTPS URLs. How to hand clips between the two.
- Music API has no duration field: steer length in the prompt
Sume's music router rejects duration and duration_seconds. Ask for a 30-second or 2-minute track in the prompt, with section timestamps, and verify.
- Retrying a Sume music job: Idempotency-Key and no double charge
Retry a timed-out Sume music request with the same Idempotency-Key and get the original job back. Lyria 3.5 varies per call, so never use a new key.
Written by Sume