FLUX 3 Action is a robot-control model, not an image model

FLUX 3 Action is an open-weight 7B world action model for robots. It predicts actions, not images, so it is no drop-in for the FLUX ids Sume serves.

4 min readSume
All posts

FLUX 3 Action is not an image or video generator you call for pictures. Black Forest Labs describes it as an open-weight 7B world-action-model for robot control, derived from the multimodal FLUX 3 backbone, that predicts actions from observations. On Sume, image requests use image model ids such as black-forest-labs/flux.2-pro.

Vendor facts are from the BFL post, read 2026-10-01. The snapshot does not show the weights license, so I do not state it here: read the license on the weights page before commercial use. Sume facts are from Image generation.

What is FLUX 3 Action?

The post says it was pretrained on "a large-scale collection of video, image and audio data" with strong emphasis on video, then adapted through joint video-action training and finetuning for a target embodiment and its action space. On the RoboLab-120 leaderboard it reports 38.3% for the single-step 7B checkpoint and 42.2% for the guidance-distilled one, against 36.8% for Cosmos 3 Nano. Those are the vendor's figures on a robotics benchmark; they say nothing about image quality.

How does that differ from an image request?

FLUX 3 Action versus a Sume image request, from the BFL post and Sume docs read 2026-10-01
QuestionFLUX 3 ActionSume image request
OutputPredicted robot actions, jointly with videoImages from POST /v1/images
InputObservations such as multi-camera robot videoA prompt plus input_references (the catalog example allows 0-10)
Where it runsYour hardware, open weightsHosted by Sume
Unknown parametersNot applicableRejected with 400 unsupported_parameter

Which FLUX ids does Sume accept for images?

The image contract lists black-forest-labs/flux.2-pro and black-forest-labs/flux.2-flex. A request that sets a parameter the selected model does not list is rejected rather than silently dropped, so sending an action-style field to an image model returns an error. For the image side of FLUX 3, see FLUX 3 image generation vs FLUX 2.

Does it matter if I only make media?

Mostly no. Unless you are building robot policies, treat this release as news about the FLUX 3 backbone. For clips, use the video catalog from GET /v1/videos/models and pick by listed duration and resolution; FLUX 3 video: 20-second clips covers that gap.

Sources

Related posts

More in Models

All Models posts

Written by Sume