MiniMax H3 on SGLang: the FSDP corruption warning and what to use
MiniMax's H3 guide warns that FSDP inference has reported data corruption; use TP plus Ulysses. The serve flags, the cross-node rule, one service per variant.

If you serve MiniMax H3 with SGLang on several GPUs, do not use FSDP for inference. MiniMax's self-hosting guide says FSDP inference has a reported data corruption issue (it links SGLang issue 34227) and tells you to prefer tensor parallelism plus Ulysses placement. In the commands the guide prints, that is --tp-size 1 with --ulysses-degree set to the GPU count.
What are the parallelism flags in the vendor commands?
Across nodes, the guide says to use --encoder-parallel replicate.
| Source | GPUs | Parallelism flags |
|---|---|---|
| README example | 4 | --num-gpus 4 --ulysses-degree 4 --performance-mode speed |
| Self-host guide, reference | 8 x B200 | --num-gpus 8 --tp-size 1 --ulysses-degree 8 --encoder-parallel auto --performance-mode speed |
What else is easy to miss?
- FL2VA and Ref2VA are separate partitions; each needs its own SGLang service instance.
- The guide lists about 108 GB of local disk for the weights.
- The guide's latency table is for 8 x B300 with Ulysses 8: 19.04 s for FL2VA and 29.12 s for Ref2VA, at roughly 83.6 to 84 GB peak VRAM per GPU in BF16.
- Do not compare quantized ComfyUI output with lossless SGLang output in benchmarks.
How do you check your own setup?
Read your launch command for any FSDP flag and remove it. Then compare against the vendor line: --tp-size 1, --ulysses-degree equal to --num-gpus, and --encoder-parallel auto on a single node. If a result looks corrupted (garbled frames or audio), re-run the same prompt with a different parallel layout before you blame the prompt.
Keep the SGLang commit pinned. The guide pins a specific commit for its reference setup.. A pinned commit means a clip made today can be reproduced next month on the same box, which matters more than the latest build.
When does a hosted job save this work?
The parallelism choice, the variant split and the pinned SGLang commit are all yours to maintain on a self-hosted node. With a hosted id the surface is POST /v1/videos with model, prompt, duration, resolution, and an optional callback_url. On Sume, minimax-h3 takes 5 to 15 seconds at 480p or 768p. Reference-to-video is the same id through input_references, so there is no second service to run.
See the video docs for the request fields. If you still self-host, keep your own notes on which flags produced which clip.
Sources
Related posts
More in Developers
- Music job timed out? Retry with the same key: one $0.125, not two
A Sume music submit that times out can be retried with the same Idempotency-Key and returns the original job. Without a key, a blind retry is a second $0.125.
- Music router 400 negative_prompt_unsupported: the fix and its cost
A non-empty negative_prompt on Sume's music router returns 400 negative_prompt_unsupported. Move the exclusions into prompt; a rejected call is not billed.
- Nano Banana 2.1 inpainting: Google's prompt template, no mask on Sume
Google edits one region of an image by prompt on Nano Banana 2.1. Sume has no mask_url for it, which is GPT Image 2.5 only. The template, and when to switch.
- Alert when a pricing_skus rate moves: a nightly diff of the model list
Save GET /v1/videos/models rates each night and diff them. A one-cent-per-second move costs $0.30 on a 30 s clip and $300 over 1,000 of them. Python, 25 lines.
Written by Sume