Remember when "AI image generator," "AI video generator," and "AI robot brain" were three separate categories with three separate pitch decks? Black Forest Labs just filed all of that under one roof, and the roof also happens to talk, sing, and occasionally install car door seals.
One Model, Four Jobs
Launched July 23, FLUX 3 is a single network handling image generation, video generation, audio, and action-prediction rather than bolting separate specialist models together. The headline feature is video: up to 20 seconds per clip with native audio — dialogue, sound effects, music — plus text-to-video, image-to-video, video-to-video, keyframe control, and multi-shot chaining. The idea is that one architecture learns spatial structure, motion, sound, and physical interaction together instead of faking the connections between them after the fact.
Access is staged: FLUX 3 Video is in early access now via an application at bfl.ai, image generation follows in the coming weeks, and Black Forest Labs says API access, private weights, and an open-weight "Dev" version are coming later this year.
The Robot Arm Is Also Watching This Video
The genuinely odd flex here is action-prediction: robotics partner mimic robotics is using FLUX 3 as the backbone for a video-action model already being tested in Audi factories, handling fiddly tasks like installing flexible door seals with reaction times around 101 milliseconds. That's a video generator moonlighting as a robot's reflexes, which is not a sentence anyone was writing two years ago.
Black Forest Labs' own head-to-head comparisons have FLUX 3 winning preference tests against Grok Imagine Video, Kling v3 Pro, Runway Gen 4.5, and Luma Ray 3.2 by wide margins — up to 93% against Luma. Take vendor-run benchmarks with the usual grain of salt, but the direction of travel — one model doing perception, generation, and control — is the more interesting story than any single win rate.
The generative-AI industry spent years arguing about image models. FLUX 3 is betting the next argument is about who can generate a whole scene, sound and all — and maybe hand it straight to a robot.
Curious how solid your own stack would hold up? Get in touch and we'll take a look.
Source: VentureBeat