ICYMI: Black Forest Labs opens FLUX 3 Video early access

FLUX 3 Video is now in early access, generating up to 20-second videos with native audio from text, images, keyframes or reference clips.

· 2 min read
BFL

Black Forest Labs has opened early access to FLUX 3 Video, the first available part of a new multimodal foundation model trained jointly on images, video, and audio. The model can create videos with native audio up to 20 seconds long in one generation, starting from text, images, keyframes, or reference clips. It can also continue video and audio, carry central elements such as a character into new scenes, produce multilingual dialogue, and chain clips into longer multi-shot sequences.

Rather than treating each medium as a separate task, FLUX 3 uses one architecture to learn how appearance, motion, and sound constrain the same event. Language connects those perceptions to instructions and goals. The system builds on Black Forest Labs' Self-Flow method for aligning multimodal generation and understanding, with training scaled across all three media types at once.

In preliminary tests, Black Forest Labs generated 10-second, 720p text-to-video clips with audio. The company says FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Seedance 2.0 and Gemini Omni Flash in 52%, Runway Gen-4.5 in 77%, and Luma Ray 3.2 in 93%. The results remain early because both the model and evaluation harness are still in development.

FLUX 3 also supports image synthesis and editing across varied styles, aspect ratios, and resolutions, with stronger handling of complex prompts and multilingual text than earlier FLUX versions in midtraining evaluations. Early access to FLUX 3 Image is due in the following weeks. Future access is planned through APIs and private weights, alongside an open-weight FLUX 3 Dev backbone for image, video, audio, and action prediction.

Black Forest Labs is positioning FLUX 3 as a step toward models that can perceive, predict, and act across physical and digital environments. Its action work follows two paths: native action prediction inside FLUX 3, and specialist models finetuned from its video backbone with limited task-specific data. The first partner project, FLUX-mimic, combines that backbone with mimic robotics' robot-learning work and is being tested on production tasks at Audi. The longer-term goal is to bring perception, action, and language prediction into one model.

Source