Black Forest Labs Releases Multimodal FLUX 3

On July 23, Black Forest Labs released FLUX 3 in early access, extending its FLUX family from image generation to a multimodal model for images, video, and audio. The company reports that FLUX 3 can generate video clips with native audio up to 20 seconds long, while an early version also underpins FLUX-mimic, a video-action model developed with mimic and tested at Audi.
Black Forest Labs released FLUX 3 in early access on July 23, introducing its first public video-generation model and expanding the FLUX family beyond still-image generation. According to the company's launch materials, the model is jointly trained on images, video, and audio within one architecture rather than combining separate modality-specific models behind a single interface.
The release is initially gated. Black Forest Labs says FLUX 3 Video and FLUX 3 Action will move through early-access phases, FLUX 3 Image will follow, and a planned FLUX 3 Dev version will provide open-weight access to a multimodal backbone. The company has not announced pricing, production service commitments, or a generally available public API.
Video, audio, and image generation
Black Forest Labs describes FLUX 3 as a multimodal foundation model intended to learn relationships among spatial structure, temporal motion, and sound. The company argues that joint training helps constrain generated outputs, for example by linking visual impacts with corresponding audio and movement over time.
The company says FLUX 3 can create videos with native audio up to 20 seconds long from text, image, or video references. Its listed capabilities include text-to-video, image-to-video, video-to-video, video-and-audio continuation, keyframe transitions, multilingual dialogue, and chaining clips into longer multi-shot sequences. Decrypt independently reported the early-access launch, the 20-second limit, and the model's planned rollout structure.
Black Forest Labs also published preliminary preference results using 10-second, 720p clips. It reports that FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, and Grok Imagine Video in up to 69%. The company labels the evaluations as early and says the model and harness remain in development. It has not published the sample size, rater count, or full methodology needed for independent reproduction.
A shared backbone for robotics
In a separate technical post, Black Forest Labs announced FLUX-mimic, a video-action model developed with robotics company mimic using an early version of FLUX 3. The companies say the system has been tested and deployed at Audi for tasks including parts kitting, component insertion, assembly, and handling flexible materials such as seals and cables.
The post does provide some hardware and latency detail: Black Forest Labs says the FLUX-mimic backbone runs in under 80 milliseconds on a single RTX 5090 GPU and that the optimized robot system reacts in 101 milliseconds. Those are company-reported figures, not independent validation. The public materials still do not disclose how many robots, tasks, or plants are involved, nor do they provide per-task success rates, intervention counts, or safety results.
For ML practitioners, the notable proposition is the attempt to use one learned representation across generative media and visual-action prediction. Video preference tests do not establish manipulation reliability. Evaluating a physical-AI deployment requires task-level success measures, failure analysis, latency under operating load, human-intervention rates, and safety evidence.
Access and open-weight status
FLUX 3 does not launch with downloadable weights or an open-source license. Black Forest Labs has announced FLUX 3 Dev as a future open-weight multimodal backbone, but it has not provided a specific release date, license, model size, or hardware requirements.
The launch gives developers a concrete new audio-video capability and a partner-backed robotics experiment. Its practical significance will depend on access, cost, reproducible evaluation, and whether the reported representation transfers beyond the initial demonstrations.
Key Points
- 1FLUX 3 extends Black Forest Labs from image generation into jointly trained image, video, and audio generation, with early access initially restricted.
- 2Black Forest Labs reports native audio-video generation up to 20 seconds, but pricing, general API availability, and reproducible evaluation details remain undisclosed.
- 3FLUX-mimic has been tested at Audi on named manipulation tasks, with company-reported RTX 5090 latency, while deployment scale, task success rates, and safety results remain unpublished.
- 4An announced open-weight FLUX 3 Dev variant could matter for researchers, but no specific release date, license, model size, or hardware requirement is available.
Scoring Rationale
FLUX 3 is a notable multimodal model release from an established image-generation lab, combining native audio-video generation with a reported robotics application. Its practitioner impact is constrained at launch by limited access, absent public API availability, no downloadable weights, and incomplete evaluation disclosure.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

