Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model that learns from images, video, and audio inside a single architecture. It is also the first FLUX model to send video, audio, and action predictions from a single set of weights.
The Black Forest Labs (BFL) research team argues that no single methodology gives a complete description of the world. Images capture spatial structure in an instant. Video restores time and highlights physical mobility. Audio reveals causal relationships between mechanical phenomena and sound. Each is treated as a harmful projection of the same underlying reality.
Training on all of them simultaneously means that the modalities inhibit each other. The sound should match the effect. Momentum must follow mass. The research team describes the FLUX 3 as its first model built entirely on that principle.
Method below: Self-flow
FLUX 3 is based on BFL’s self-flow methodology to align multimodal generation and understanding into a single architecture. Self-flow combines the flow matching objective with the self-supervised feature reconstruction objective. The reference implementation on GitHub is Apache-2.0 and uses SiT-XL/2 with per-token timestep conditioning. It trains with a 25% per-token mask ratio and self-distillation from an EMA teacher at layer 20 to a student at layer 8.
The checkpoint released is an ImageNet 256×256 research model, not FLUX 3. BFL says it has ‘significantly increased compute and data resources’ on a single approach to train FLUX 3 simultaneously on video, images and audio. Self-Flow was only introduced in March 2026, so this launch is nothing new. What is new is scale.
what does flux 3 video
FLUX 3 Video produces clips up to 20 seconds long with native audio in a single generation. Supported modes cover text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video for controlled transitions, and generative video-audio continuity from input video and audio.
BFL also lists strong typography generation with multilingual dialogue, agentic chaining of clips in multi-shot sequences, and animated design. The BFL team reports particular strengths in associating sounds with human facial expressions and physical events.
Display
The BFL team publishes preliminary human preference results. The setup was a 10-second text-to-video clip at 720p with audio. Flux 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77% of comparisons. This figure is up to 69% compared to Grok Imagine Video, then Cling V3 Pro at 60%, Happy Horse V1 at 59% and Happy Horse 1.1 at 57%. Compared to the Sedan 2.0 and Gemini Omni Flash, the result is 52%, which is close to a coin toss.
interactive Explorer
enter
key takeaways
- FLUX 3 is a flow matching backbone jointly trained on image, video and audio.
- FLUX 3 video is generated up to 20 seconds with native audio in one generation.
- Video prediction consumes more than 95% of the training computation; Audio is less than 0.5% token.
- The same backbone powers FLUX-mimic, a robot policy that runs in less than 80 ms on an RTX 5090.
- Access is gated: video and action are in early access, images come next, open weights come last.
Check out the Flux 3 announcement, Flux 3 x copy technical post and self-flow paper. All credit for this research goes to the researchers of this project.

Michael Sutter is a data science professional and holds a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michael excels in transforming complex datasets into actionable insights.