Black Forest Labs has officially unveiled FLUX 3, a multimodal foundation model that integrates image, video, audio, and robot action prediction into a single, unified architecture. By training on these modalities simultaneously, the model ensures that generated outputs maintain physical consistency, where motion obeys mass and sound aligns precisely with mechanical events. This release marks the first time a FLUX model has delivered video, audio, and action prediction from a single set of weights.
The Self-Flow Architecture
FLUX 3 is built upon the Self-Flow methodology, which aligns multimodal generation and understanding within one framework. This approach combines the flow matching objective with a self-supervised feature reconstruction objective. While the underlying method was introduced in March 2026, the launch of FLUX 3 represents a significant scaling up of compute and data resources. The model utilizes per-token timestep conditioning and employs a training process involving self-distillation from an EMA teacher to a student model.
Capabilities in Video and Audio Generation
FLUX 3 Video is capable of generating clips up to 20 seconds long in a single pass, complete with native audio. The model supports a wide array of generation modes, including text-to-video, image-to-video, video-to-video, and keyframe-to-video for controlled transitions. Additionally, the system excels at generative video-audio continuation, multilingual dialogue, and the creation of animated typography. According to the research team, the model demonstrates particular proficiency in rendering human facial expressions and associating specific sounds with physical actions.
Performance and Robot Integration
In preliminary human preference evaluations, FLUX 3 demonstrated strong performance against existing industry models. When compared to Luma Ray 3.2, FLUX 3 was preferred in 93% of 10-second text-to-video comparisons, and it outperformed Runway Gen-4.5 in 77% of cases. The model also showed competitive results against other industry benchmarks, including Kling v3 Pro, Grok Imagine Video, and Gemini Omni Flash.
Beyond media generation, the same backbone powers FLUX-mimic, a robot policy capable of running in under 80 milliseconds on a single RTX 5090. While video prediction consumes more than 95% of the training compute, audio processing accounts for less than 0.5% of tokens. Black Forest Labs has implemented a phased release strategy, with video and action capabilities currently in early access, followed by image generation, and finally, the release of open weights.

Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!