MiniMax H3 FL2VA vs Ref2VA: reference requests need own server
MiniMax splits H3 into FL2VA (text, first/last frame) and Ref2VA (images, video, audio references) checkpoints. Which tasks go where, and the Sume model ids.

MiniMax ships H3 as two checkpoint partitions for self-hosting. FL2VA handles text-to-video with audio and first/last-frame image-to-video; Ref2VA handles reference-driven generation with images, video and audio. The guide warns not to send ref2va requests to a server started for FL2VA, because they need different weights.
That comes from the Run and Self-host MiniMax H3 guide, read on 2026-10-02. On a hosted API you never see the split; on Sume you pick a model and send frame_images or input_references as described in Video generation.
Which task goes to which partition?
The table below is MiniMax's own mapping. Mixing them up is a deployment error, not a prompt error: a request for the wrong task simply has no matching weights on the server.
| Partition | Task values | Use | Note |
|---|---|---|---|
| FL2VA | t2va, fl2va | Text-to-video, first/last-frame image-to-video | Default for the SGLang server |
| Ref2VA | ref2va | Images, video and audio references | Separate service on a different port |
What does that mean for how you design a service?
If you run both, you need two processes (or one framework that loads both). The guide says SGLang needs a separate service on another port for Ref2VA, while the experimental vLLM-Omni route can load both partitions in one service at about 270 GB of storage. It names a reference config of 2 x RTX 5090 with layerwise offload and over 200 GB of system RAM for that route.
So a gateway in front must route by task, not by model name alone. A simple rule is that any request carrying reference media goes to the Ref2VA port, and everything else, including a first or last frame, goes to FL2VA.
def target(task: str) -> str:
return {
"t2va": "http://127.0.0.1:30010",
"fl2va": "http://127.0.0.1:30010",
"ref2va": "http://127.0.0.1:30011",
}[task]How does this show up on Sume?
It does not. On Sume, minimax-h3 and minimax-h3-max take a single request shape. frame_images entries with a frame_type of first_frame or last_frame make an image-to-video request, and input_references make a reference-to-video request. If both are present, frame_images takes precedence and the request is treated as image-to-video.
The docs also say audio and video references are honored by MiniMax H3 and H3 Max, and that Sume reports the supported reference types per model in supported_input_references. Check that field before sending, because a type a model does not list is not accepted.
What should you check before choosing?
Decide the tasks you need. If you only need text and frame control, one FL2VA server is enough. If you need references, budget for a second service or the heavier single-service route, and read the license first: the same guide lists the US, EU, UK and South Korea as Excluded Territories.
If you would rather not run either, send the request to a hosted model and keep one code path.
Sources
Related posts
More in Models
- MiniMax H3 Max Recast API: swap people in a video, fal price vs Sume
H3 Max Recast swaps people in a source video for reference photos, keeping motion, cuts and audio. fal lists $0.30 a second at 768p; what Sume accepts.
- MiniMax H3 sound design prompts: direct the audio like the picture
fal's H3 guide says to direct audio as deliberately as picture: name sonic elements, not 'music'. What it looks like in a Sume request, and what you can't set.
- MiniMax-Hailuo-2.3-Fast: an image-to-video model id, no Sume twin
MiniMax lists MiniMax-Hailuo-2.3-Fast in its image-to-video API but not in its text-to-video enum. What that means, and what Sume offers for first frames.
- MiniMax Hailuo 2.3 vs H3 API: model ids, durations, prompt limits
MiniMax's API still lists MiniMax-Hailuo-2.3 (6 or 10 s, 2,000-character prompt) beside MiniMax-H3 (4-15 s, 7,000). Which one Sume lists, and why it matters.
Written by Sume