Black Forest Labs announced FLUX 3 on Thursday, and the news is not that an image-generation company now does video. It is that the same model also predicts robot actions, and that almost none of it has open weights. The Freiburg company describes FLUX 3 as a single multimodal foundation model trained jointly on images, video and audio, arriving as four things: FLUX 3 Video in early access now, FLUX 3 Action with selected robotics partners now, FLUX 3 Image in the coming weeks, and FLUX 3 Dev, the open-weight backbone, at some point later in 2026.
The announcement reached Hacker News overnight and was on the front page inside a couple of hours, which is not surprising. BFL is the company whose openly released FLUX.1 weights became the default image model for a large share of independent tooling. This launch runs in the opposite direction.
RelatedMistral bets on a 'fat but sparse' open-weight MoE
- One backbone, four doors. Video means API access plus private weights for initial partners. Action is partners only. Image is not out yet. The open piece, FLUX 3 Dev, has no date beyond "later in 2026."
- Clips run up to 20 seconds with native audio, evaluated at 720p, covering text-to-video, image-to-video, video-to-video and keyframe-to-video from one model.
- Every benchmark figure is BFL's own preference test. Two of them sit at 52 percent, which against Seedance 2.0 and Gemini Omni Flash is a coin flip rather than a win.
- FLUX-mimic fine-tunes on 30 minutes of robot data where the same result previously needed 30 hours or more. That ratio, not the video reel, is the real argument.
What actually shipped, and what did not?
FLUX 3 Video is the only piece an outside customer can touch today, and even that is early access: an API plus private weights handed to initial partners. Black Forest Labs names Canva, Burda, Magnific (formerly Freepik), Krea and Picsart among the companies it works with, and its models already sit inside Adobe Photoshop. FLUX 3 Action, the robot-control head, went to what the company calls selected research and commercial robotics partners. FLUX 3 Image is due "in the following weeks." No pricing was published. Neither was a parameter count, which has become normal for frontier releases and is still worth flagging when the whole pitch is a unified architecture.
The video specifics are concrete: up to 20 seconds per generation with audio produced by the same model rather than dubbed on afterwards, tested at 720p, plus video continuation, keyframe transitions, multilingual dialogue, on-screen typography, and clip chaining for longer sequences. Text rendering, historically the weakest part of every FLUX generation, is called out as significantly improved.
Why would an image company build a robotics model?
Because BFL's thesis is that you cannot extract physical understanding from still frames. CEO Robin Rombach, one of the people behind Stable Diffusion before he co-founded BFL, put it bluntly in the announcement: "You can't cheat reality. A model that only learns images can only generate images." The company's framing is that training on images, video and audio together, under an approach it calls Self-Flow, forces the model to learn spatial structure, motion, sound and physical interaction as one thing instead of four separate tasks. Rombach's other line reads as the mission statement: "True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results."
The proof point is FLUX-mimic, built with the robotics company mimic. Its CTO Elvis Nava names the constraint that makes this worth attention: "The hardest part of robotics is data." Teleoperation data is slow and expensive to collect, which is why robot learning stayed bottlenecked for years while language and image models raced ahead. BFL's claim is that fine-tuning FLUX-mimic on 30 minutes of robot data reaches what previously took more than 30 hours, because the video backbone already encodes how objects move and collide. Audi's Production Lab is named as a test site, with Christoph Schneider from that team quoted on pushing the frontier of physical AI. A 60x reduction in data requirements, if it survives contact with other labs' hardware, matters far more than another text-to-video model does.
How good is FLUX 3 Video, actually?
Read the numbers carefully, because BFL ran them itself. These are human preference comparisons, vendor-reported, not an independent benchmark, and the spread inside them tells a clearer story than any single figure.
Against the second tier the wins are real: 93 percent preference over Luma Ray 3.2 and 77 percent over Runway Gen-4.5 are not close calls. Against the current leaders, the same tests report 52 percent versus Seedance 2.0 and 52 percent versus Gemini Omni Flash. At that margin, on a preference test, the honest word is parity. Kling v3 Pro at 60 percent is a modest edge. Credit where it is due: BFL published the 52s rather than quietly cropping the chart, which is more than several competitors have managed. But the correct reading of its own data is that FLUX 3 Video arrives at the frontier instead of ahead of it, and that video quality is converging fast enough for a first-generation entrant to land level with incumbents.
| FLUX 3 Video | FLUX 3 Image | FLUX 3 Action | FLUX 3 Dev | |
|---|---|---|---|---|
| Output | Clips to 20s, native audio | Still images | Robot action prediction | Multimodal backbone |
| Available | Early access now | Coming weeks | Now, restricted | Later in 2026 |
| Weights | Private, partners only | Not stated | Private | Open |
| Who gets access | API plus initial partners | Not stated | Selected research and robotics partners | Anyone, once released |
Is the open-weights company going closed?
This is the part the FLUX community actually cares about. BFL built its reputation and its distribution on openly released weights, and says models from its founders have passed half a billion downloads. FLUX 3 inverts the sequence: private weights first, partner access, open release deferred to "later in 2026" and then only for FLUX 3 Dev, the backbone, not the flagship video model. If your work depends on running weights locally, nothing usable shipped on Thursday.
There is a defensible reason for the ordering. A model that also predicts robot actions carries different risk than an image generator, and 20-second video inference with audio is expensive enough that an API is the only practical way to control cost early on. It is still a real change in posture, and the ecosystem that made FLUX ubiquitous now waits in line behind Canva and Audi.
RelatedLaguna S 2.1: 8B Active Params, 70% Terminal-Bench
- Aug 2024FLUX.1 ships with open weights becomes the default open image model for independent tooling
- Jul 21, 2026A FLUX 3 placeholder page appears on bfl.ai, then is pulled two days of speculation follow
- Jul 23, 2026FLUX 3 announced from Freiburg Video and Action in early access, Image pending, no pricing
- Coming weeksFLUX 3 Image early access
- Later 2026FLUX 3 Dev open weights the piece the open-source ecosystem is waiting on
What does it mean for the market?
Black Forest Labs is valued at $3.25 billion having raised more than $450 million, with a16z, Salesforce Ventures, NVIDIA, Adobe Ventures, Figma Ventures and Canva on the cap table. That investor list is itself a signal: NVIDIA sells the compute, Adobe ships FLUX inside Photoshop, Canva is both investor and customer. The strategic value here is priced in by the companies positioned to distribute it.
For anyone watching the public side, the exposure is indirect but readable. A fourth credible video model entering a market already served by Runway, Luma, Kuaishou's Kling and ByteDance's Seedance puts downward pressure on per-second generation pricing, which is bad for pure-play video-gen startups and roughly neutral for platforms that resell generation as a feature. Adobe (ADBE) is the clearest listed name with a stake in this going well, since its generative features increasingly lean on partner models. NVIDIA (NVDA) benefits either way. The signal for investors is not in the video benchmarks: it is whether action prediction converts into paying robotics deployments, because that is the part of this story that would justify the valuation on something other than creative tooling. This is analysis, not investment advice.
Who should care today?
If you ship a creative product on an API, FLUX 3 Video is worth joining early access for, with the caveat that you cannot model its cost yet. If you build on open weights, your date is the FLUX 3 Dev release, and you should plan as though it slips. If you work in robot learning, the FLUX-mimic data-efficiency claim is the most testable thing in the whole announcement, and the first place independent work will either confirm the story or puncture it.
- Whether FLUX 3 Dev actually lands. "Later in 2026" is the vaguest date in the announcement, and it is the one the open ecosystem is planning around.
- Independent video evaluations. Every number published so far is BFL's own. The 52 percent results against Seedance 2.0 and Gemini Omni Flash need a neutral harness before anyone treats them as settled.
- The 30-minute robot-data claim on someone else's hardware. Reproduced outside mimic and Audi, it reframes robot learning. Unreproduced, it is a demo.
- Pricing. No per-second cost was published. Twenty seconds of video with generated audio is not cheap to serve, and the number will decide who can build on it.
Our take
The video model is the announcement; the action model is the strategy. Video generation is on a fast path to becoming a commodity feature, and BFL's own preference numbers say as much: level with the leaders, not ahead of them. What is genuinely new is a company arguing that a generative video backbone is the cheapest available source of physical world knowledge, then shipping a robotics model to prove the point. That bet deserves to be taken seriously, and it explains the closed launch better than the usual safety framing does. The cost is credibility with the open-source community that made FLUX a standard, and BFL has until "later in 2026" to pay it back.
- OfficialBlack Forest Labs: FLUX 3 model variants, 20s clips, 720p evals, preference-test percentages
- Press releaseGlobeNewswire: Black Forest Labs unveils FLUX 3 $3.25B valuation, $450M raised, partner and executive quotes
- ReferenceGenZTech: How AI image generators work background on the diffusion pipeline FLUX builds on
Original analysis by GenZTech. Announcement: Black Forest Labs, FLUX 3.
