Black Forest Labs Launches FLUX 3 Action, a 7B Robot Control Model Leading RoboLab-120

Black Forest Labs has introduced FLUX 3 Action, a 7B open-weights model for robot control that combines video prediction with action generation. Given camera frames, robot state, and a text instruction, it forecasts future video frames and the next action chunk. It leads RoboLab-120 with 42.92% task success, and can run on 24 GB GPUs with FP8 quantization and text-encoder offload under a non-commercial license.
Black Forest Labs, known for FLUX image generators, has extended that line into robotics with a 7B open-weights controller. The system ingests camera views, robot state, and a language command, then jointly forecasts upcoming video and an action segment. It currently sits atop RoboLab-120 at 42.92% success.
Its backbone comes from multimodal FLUX 3 pretraining, where video supplied more than 95% of tokens. Midtraining blended pretraining samples with video paired to actions, covering 14 embodiments and a shared EE50 end-effector space. Deployment needs roughly 32 GB in BF16, or 24 GB cards via FP8 and moving the text encoder off the GPU, under a non-commercial license.
The release could lower the compute barrier for robotics research, letting smaller labs test advanced control policies on 24 GB GPUs. It may speed experimentation in physical AI, but the non-commercial license could limit startups and product teams. Workers and users might eventually see more capable, cheaper robots, while safety and reliability remain open questions.