Black Forest Labs has opened early access to FLUX 3 Video, the first available part of a new multimodal foundation model trained jointly on images, video, and audio. The model can create videos with native audio up to 20 seconds long in one generation, starting from text, images, keyframes, or reference clips. It can also continue video and audio, carry central elements such as a character into new scenes, produce multilingual dialogue, and chain clips into longer multi-shot sequences.
Introducing FLUX 3.
— Black Forest Labs (@bfl_ai) July 23, 2026
One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.
FLUX 3 Video is now available in early access (link below).
Jointly trained in one unified architecture, our model can be extended to… pic.twitter.com/voQ5iUJJZY
Rather than treating each medium as a separate task, FLUX 3 uses one architecture to learn how appearance, motion, and sound constrain the same event. Language connects those perceptions to instructions and goals. The system builds on Black Forest Labs' Self-Flow method for aligning multimodal generation and understanding, with training scaled across all three media types at once.
In preliminary tests, Black Forest Labs generated 10-second, 720p text-to-video clips with audio. The company says FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Seedance 2.0 and Gemini Omni Flash in 52%, Runway Gen-4.5 in 77%, and Luma Ray 3.2 in 93%. The results remain early because both the model and evaluation harness are still in development.
FLUX 3 also supports image synthesis and editing across varied styles, aspect ratios, and resolutions, with stronger handling of complex prompts and multilingual text than earlier FLUX versions in midtraining evaluations. Early access to FLUX 3 Image is due in the following weeks. Future access is planned through APIs and private weights, alongside an open-weight FLUX 3 Dev backbone for image, video, audio, and action prediction.
Black Forest Labs is positioning FLUX 3 as a step toward models that can perceive, predict, and act across physical and digital environments. Its action work follows two paths: native action prediction inside FLUX 3, and specialist models finetuned from its video backbone with limited task-specific data. The first partner project, FLUX-mimic, combines that backbone with mimic robotics' robot-learning work and is being tested on production tasks at Audi. The longer-term goal is to bring perception, action, and language prediction into one model.