Black Forest Labs, the German AI lab behind the FLUX image models, launched FLUX 3 in early access on July 23 — its first system trained to generate images, video, and audio within a single model rather than stitching together separate tools.
What’s new
Previous FLUX releases handled still images only. FLUX 3 Video generates clips up to 20 seconds long with synchronized audio — dialogue, sound effects, and ambient noise matched to the visuals — from text prompts, existing images, or other video, and supports multilingual dialogue and extending a clip from a fixed keyframe. A companion image-generation mode, FLUX 3 Image, is due to open to early access “in the coming weeks,” the company said, and an open-weight FLUX 3 Dev release is planned for later this year alongside API access.
Black Forest Labs says internal human-preference testing found reviewers favored FLUX 3 video clips over rival models Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%. The company has not published its testing methodology, and the figures have not been independently verified.
A push into robotics
FLUX 3’s video backbone doubles as a foundation for robotics: a new component called FLUX 3 Action, built with robotics startup mimic robotics, predicts physical motion the same way the model predicts video frames, treating a robot arm’s next move like the next frame in a clip. Black Forest Labs said the approach can be fine-tuned for a new manipulation task with as little as 30 minutes of robot data, and that automaker Audi is testing it for production-line work — a use case similar in spirit to the action-prediction approach behind Vision-Language-Action models already deployed in humanoid robots.
Founded in 2024 by former Stability AI researchers Robin Rombach, Andreas Blattmann, and Patrick Esser, Black Forest Labs raised $300 million in a Series B round in December 2025 at a $3.25 billion valuation. Pricing for FLUX 3 has not been disclosed; access is currently limited to companies that apply for early access.