FLUX 3

Black Forest Labs

FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.

Specs

Model

Family
FLUX
Variant
3
Weights
Proprietary

Modalities

Input
Text Image Video Audio
Output
Video Audio

Dates

Announced
23 Jul 2026
Released
4 Aug 2026
Black Forest Labs

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence

FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.

bfl.ai

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.