MiniMax H3 FL2VA vs Ref2VA: reference requests need own server

MiniMax splits H3 into FL2VA (text, first/last frame) and Ref2VA (images, video, audio references) checkpoints. Which tasks go where, and the Sume model ids.

5 min readSume
All posts

MiniMax ships H3 as two checkpoint partitions for self-hosting. FL2VA handles text-to-video with audio and first/last-frame image-to-video; Ref2VA handles reference-driven generation with images, video and audio. The guide warns not to send ref2va requests to a server started for FL2VA, because they need different weights.

That comes from the Run and Self-host MiniMax H3 guide, read on 2026-10-02. On a hosted API you never see the split; on Sume you pick a model and send frame_images or input_references as described in Video generation.

Which task goes to which partition?

The table below is MiniMax's own mapping. Mixing them up is a deployment error, not a prompt error: a request for the wrong task simply has no matching weights on the server.

H3 checkpoint partitions, read 2026-10-02
PartitionTask valuesUseNote
FL2VAt2va, fl2vaText-to-video, first/last-frame image-to-videoDefault for the SGLang server
Ref2VAref2vaImages, video and audio referencesSeparate service on a different port

What does that mean for how you design a service?

If you run both, you need two processes (or one framework that loads both). The guide says SGLang needs a separate service on another port for Ref2VA, while the experimental vLLM-Omni route can load both partitions in one service at about 270 GB of storage. It names a reference config of 2 x RTX 5090 with layerwise offload and over 200 GB of system RAM for that route.

So a gateway in front must route by task, not by model name alone. A simple rule is that any request carrying reference media goes to the Ref2VA port, and everything else, including a first or last frame, goes to FL2VA.

def target(task: str) -> str:
    return {
        "t2va": "http://127.0.0.1:30010",
        "fl2va": "http://127.0.0.1:30010",
        "ref2va": "http://127.0.0.1:30011",
    }[task]

How does this show up on Sume?

It does not. On Sume, minimax-h3 and minimax-h3-max take a single request shape. frame_images entries with a frame_type of first_frame or last_frame make an image-to-video request, and input_references make a reference-to-video request. If both are present, frame_images takes precedence and the request is treated as image-to-video.

The docs also say audio and video references are honored by MiniMax H3 and H3 Max, and that Sume reports the supported reference types per model in supported_input_references. Check that field before sending, because a type a model does not list is not accepted.

What should you check before choosing?

Decide the tasks you need. If you only need text and frame control, one FL2VA server is enough. If you need references, budget for a second service or the heavier single-service route, and read the license first: the same guide lists the US, EU, UK and South Korea as Excluded Territories.

If you would rather not run either, send the request to a hosted model and keep one code path.

Sources

Related posts

More in Models

All Models posts

Written by Sume