FLUX 3 x mimic: one backbone for video, audio and robot actions

BFL says FLUX 3 is one model trained jointly on images, video and audio, and FLUX-mimic adds robot actions. What that does and does not tell a media API user.

4 min readSume
All posts

FLUX-mimic is a video-action model for robots, built by Black Forest Labs and mimic robotics on an early FLUX 3 backbone. The post states that FLUX 3 is "one model, jointly trained across images, video and audio from the beginning". That describes how the model was trained; it does not tell you which inputs a hosted API accepts, which is listed per model.

Vendor claims are from the BFL post dated July 23, 2026, read 2026-10-01; they are the vendor's statements, which I did not test. Sume facts are from Video generation and Image generation.

What does the post claim about the backbone?

It says video prediction is the most demanding part of training, "over 95% of the total compute costs", and that audio is less than 0.5% of the tokens in a 720p video with audio. When action prediction was added to a large training run, human ratings on text-to-video and image-to-video "initially fell by up to 10%", and after 3500 steps the model regained its previous quality while also predicting actions. The conclusion drawn is that video generation and action prediction can share one foundation.

What does joint training mean for a hosted catalog?

Very little on its own. A hosted catalog exposes a fixed set of inputs per model id, and requests are checked against it.

Training claim versus catalog behavior, BFL post and Sume docs read 2026-10-01
TopicBFL postSume catalog
ModalitiesImages, video and audio trained jointlyEach video model lists supported_input_references, for example image_url, video_url, audio_url
Clip lengthNot stated as an API limitPer model; most catalog models top out at 15 seconds
Image inputsNot stated as an API limitCatalog example allows input_references from 0 to 10
Unlisted parameterNot applicableRejected with 400 unsupported_parameter

Can I run FLUX-mimic through an API?

The part of the post I read describes a deployed robotics system, not a public media endpoint, so do not assume one. If you want video from Sume, read GET /v1/videos/models for current ids, durations and resolutions; FLUX 3 video: 20-second clips covers what is on offer instead.

What is the practical takeaway?

Treat the post as research news about one backbone doing two jobs. For media work, pick by what the model lists, not by what it was trained on. The companion release is covered in FLUX 3 Action is a robot-control model.

Sources

Related posts

More in Models

All Models posts

Written by Sume