FLUX 3 Action is a robot-control model, not an image model
FLUX 3 Action is an open-weight 7B world action model for robots. It predicts actions, not images, so it is no drop-in for the FLUX ids Sume serves.

FLUX 3 Action is not an image or video generator you call for pictures. Black Forest Labs describes it as an open-weight 7B world-action-model for robot control, derived from the multimodal FLUX 3 backbone, that predicts actions from observations. On Sume, image requests use image model ids such as black-forest-labs/flux.2-pro.
Vendor facts are from the BFL post, read 2026-10-01. The snapshot does not show the weights license, so I do not state it here: read the license on the weights page before commercial use. Sume facts are from Image generation.
What is FLUX 3 Action?
The post says it was pretrained on "a large-scale collection of video, image and audio data" with strong emphasis on video, then adapted through joint video-action training and finetuning for a target embodiment and its action space. On the RoboLab-120 leaderboard it reports 38.3% for the single-step 7B checkpoint and 42.2% for the guidance-distilled one, against 36.8% for Cosmos 3 Nano. Those are the vendor's figures on a robotics benchmark; they say nothing about image quality.
How does that differ from an image request?
| Question | FLUX 3 Action | Sume image request |
|---|---|---|
| Output | Predicted robot actions, jointly with video | Images from POST /v1/images |
| Input | Observations such as multi-camera robot video | A prompt plus input_references (the catalog example allows 0-10) |
| Where it runs | Your hardware, open weights | Hosted by Sume |
| Unknown parameters | Not applicable | Rejected with 400 unsupported_parameter |
Which FLUX ids does Sume accept for images?
The image contract lists black-forest-labs/flux.2-pro and black-forest-labs/flux.2-flex. A request that sets a parameter the selected model does not list is rejected rather than silently dropped, so sending an action-style field to an image model returns an error. For the image side of FLUX 3, see FLUX 3 image generation vs FLUX 2.
Does it matter if I only make media?
Mostly no. Unless you are building robot policies, treat this release as news about the FLUX 3 backbone. For clips, use the video catalog from GET /v1/videos/models and pick by listed duration and resolution; FLUX 3 video: 20-second clips covers that gap.
Sources
Related posts
More in Models
- FLUX 3 Image bounding boxes vs Sume masked edit: what each takes
FLUX 3 Image takes a JSON list of boxes at the end of the prompt. Sume's image edit takes a mask image: mask_image_url with image_urls. Differences, no overlap.
- FLUX 3 Image 10 reference images vs Sume's 1-10 image_urls
FLUX 3 Image takes up to ten references named by prompt position. Sume's image edit takes one to ten public image_urls, with no FLUX 3 id in the docs.
- FLUX 3 x mimic: one backbone for video, audio and robot actions
BFL says FLUX 3 is one model trained jointly on images, video and audio, and FLUX-mimic adds robot actions. What that does and does not tell a media API user.
- FLUX 3 video continuation caps at 15 s: extending a clip on Sume
BFL cut FLUX 3 video continuation (v2v) to 15 seconds on Aug 17; other modes stay at 20. Sume has no continuation mode: chain a last frame instead.
Written by Sume