FLUX 3 x mimic: one backbone for video, audio and robot actions
BFL says FLUX 3 is one model trained jointly on images, video and audio, and FLUX-mimic adds robot actions. What that does and does not tell a media API user.

FLUX-mimic is a video-action model for robots, built by Black Forest Labs and mimic robotics on an early FLUX 3 backbone. The post states that FLUX 3 is "one model, jointly trained across images, video and audio from the beginning". That describes how the model was trained; it does not tell you which inputs a hosted API accepts, which is listed per model.
Vendor claims are from the BFL post dated July 23, 2026, read 2026-10-01; they are the vendor's statements, which I did not test. Sume facts are from Video generation and Image generation.
What does the post claim about the backbone?
It says video prediction is the most demanding part of training, "over 95% of the total compute costs", and that audio is less than 0.5% of the tokens in a 720p video with audio. When action prediction was added to a large training run, human ratings on text-to-video and image-to-video "initially fell by up to 10%", and after 3500 steps the model regained its previous quality while also predicting actions. The conclusion drawn is that video generation and action prediction can share one foundation.
What does joint training mean for a hosted catalog?
Very little on its own. A hosted catalog exposes a fixed set of inputs per model id, and requests are checked against it.
| Topic | BFL post | Sume catalog |
|---|---|---|
| Modalities | Images, video and audio trained jointly | Each video model lists supported_input_references, for example image_url, video_url, audio_url |
| Clip length | Not stated as an API limit | Per model; most catalog models top out at 15 seconds |
| Image inputs | Not stated as an API limit | Catalog example allows input_references from 0 to 10 |
| Unlisted parameter | Not applicable | Rejected with 400 unsupported_parameter |
Can I run FLUX-mimic through an API?
The part of the post I read describes a deployed robotics system, not a public media endpoint, so do not assume one. If you want video from Sume, read GET /v1/videos/models for current ids, durations and resolutions; FLUX 3 video: 20-second clips covers what is on offer instead.
What is the practical takeaway?
Treat the post as research news about one backbone doing two jobs. For media work, pick by what the model lists, not by what it was trained on. The companion release is covered in FLUX 3 Action is a robot-control model.
Sources
Related posts
More in Models
- FLUX 3 video continuation caps at 15 s: extending a clip on Sume
BFL cut FLUX 3 video continuation (v2v) to 15 seconds on Aug 17; other modes stay at 20. Sume has no continuation mode: chain a last frame instead.
- FLUX 3 Video 4K uhd (3840x2176) and which Sume video models do 4K
FLUX 3 Video's uhd resolution is 3840 x 2176 at 16:9. On Sume, a 4K request goes to a model whose supported_resolutions lists it; here is what the docs say.
- Flux TTS expressivity -2 to 2 vs Sume's emotion guide
Deepgram's Flux TTS expressivity runs -2 to 2 (0 nominal). Sume TTS has no such dial: generation_config takes volume, speed and a free-text emotion guide.
- Gemini 3.1 flash image preview shutdown: Sume model ids
Google lists gemini-3.1-flash-image-preview and gemini-3-pro-image-preview for June 25, 2026 shutdown. Sume callers use catalog ids like google/nano-banana-2.
Written by Sume