FLUX 3 generates image, video, and audio from one model — Black Forest Labs' bet on natively multimodal generation
TL;DR: Black Forest Labs — the lab behind the FLUX image models — unveiled FLUX 3 on July 23, calling it the first natively multimodal architecture: image, video, and audio from a single set of jointly-trained weights, not separate models stitched together. FLUX 3 Video makes clips up to 20 seconds with synchronized native audio (dialogue, SFX, ambient), plus text/image/video-to-video and keyframe transitions. In Black Forest Labs’ own evals, reviewers preferred it over Runway Gen-4.5 (77%) and Luma Ray 3.2 (93%) — vendor numbers, not independent, and notably Veo — the incumbent at video-with-native-audio, exactly FLUX 3’s headline feature — wasn’t in the comparison. What ships today: Video and a robotics-focused Action model in early access; image generation “in the coming weeks,” open weights later in 2026. What this means for you: a genuinely new architecture for AI video with native sound — promising, but verify the claims when you get access, because most of it isn’t broadly available yet.
What was announced
On July 23, 2026, Black Forest Labs unveiled FLUX 3, describing it as “the first natively multimodal architecture, bringing together image, video, audio and action generation within a single model.” The architectural claim is the story: rather than assembling separate image, video, and audio models behind a shared interface — the industry norm — FLUX 3 is trained jointly across all three modalities and generates them from one set of weights.
What that produces, per the company and VentureBeat:
- FLUX 3 Video — clips up to 20 seconds with audio generated alongside the visuals: dialogue, sound effects, and ambient noise. Supports text-to-video, image-to-video, and video-to-video, plus keyframe-controlled transitions and multilingual dialogue.
- FLUX 3 Action — aimed at robotics, not creators; manufacturing partners (including Audi, via robotics firm mimic) are testing it for complex manipulation tasks.
- Claimed quality — in Black Forest Labs’ own evaluations, human reviewers preferred FLUX 3 over Runway Gen-4.5 in 77% of head-to-heads and over Luma Ray 3.2 in 93%.
What’s actually available: only Video and Action are in early access now (APIs and select partners). Image generation follows “in the coming weeks,” and the open-weight “Dev” version isn’t due until later in 2026. So the “one model for everything” headline is the architecture; the shipping products are Video and the robotics model.
Why this matters
1. Native multimodality is a real architectural shift, if it delivers. Today’s AI-media stacks glue an image model to a separate video model to a separate audio model. That’s why generated video so often has sound that doesn’t quite fit, or a look that drifts between the still and the motion. Training one model jointly on image, video, and audio targets exactly that seam — coherence across modalities by design, especially synchronized sound. Native audio in particular has been the missing piece in most AI video (Veo being the notable exception); a model that generates picture and sound together rather than dubbing after the fact is a genuine step, not a spec bump. The caveat is “if it delivers” — architecture claims are easy, coherent 20-second clips with convincing synced audio are hard.
2. The competitive framing targets Runway and Luma — but skips the biggest names. Black Forest Labs benchmarked against Runway Gen-4.5 and Luma Ray 3.2, and beat both handily in its own evals. Two honest asterisks: these are vendor-run preference tests (the lab chose the prompts, the comparisons, and what to publish), and Google’s Veo — the incumbent leader at video-with-native-audio, which is FLUX 3’s own headline feature — wasn’t in the comparison. That omission is the informative one: you benchmark against the rivals you can beat, and skipping the model that already does the exact thing you’re claiming as new is conspicuous. It doesn’t mean FLUX 3 is bad — it may be excellent — but “beats Runway and Luma” is a narrower claim than “best AI video model,” and the gap between those two statements is exactly where buyers get misled.
3. Black Forest Labs has credibility, which raises the stakes. This isn’t a random startup. BFL’s FLUX image models became a genuine open-weight standard for image generation, competing with Midjourney and Stable Diffusion. A lab with that track record moving into jointly-trained video+audio is worth taking seriously — and the promised open-weight Dev release later in 2026 matters most of all. If a high-quality, natively-multimodal model ships with open weights, it does to AI video what FLUX did to AI image and what DeepSeek did to text: pressure the closed leaders on price and access. That’s the development to watch, more than the launch-day evals.
4. The robotics angle is a tell about where “video models” are heading. FLUX 3 Action — a generation model pointed at robot manipulation rather than content — signals that “video generation” and “world modeling for robotics” are converging. A model that understands how objects move, sound, and interact well enough to make a convincing 20-second clip is learning something a robot can use. Most creators won’t care about Action directly, but it explains why a media lab is suddenly training on “action”: the same physical-world understanding powers both. It’s the clearest sign yet that generative video and embodied AI are becoming the same research problem.
5. “Limited release” is the pattern, and it shapes what you can trust. Like GPT-5.6’s staged rollout and Gemini 3.5 Pro’s endless preview, FLUX 3 launched as an announcement plus partial early access, not a product you can fully use. That’s now standard, and it means the marketing lands weeks before independent verification is possible. Read launch-day claims — including the 77%/93% numbers — as hypotheses until creators outside the partner program put it through real work.
What this means for you
- If you make AI video: this is worth getting on the waitlist for, specifically for native synced audio and 20-second length. But don’t switch pipelines on the announcement — test FLUX 3 against your current tool (Runway, Veo, Kling) on your actual prompts once you have access.
- Weight the vendor evals accordingly. “Beats Runway and Luma” is Black Forest Labs’ own testing, and it excludes Veo and Sora. Useful signal, not a ranking. Compare options in the best AI video tools guide.
- If you generate images: nothing for you yet — FLUX 3 image is “coming weeks.” For now, FLUX’s existing models, Midjourney, and Nano Banana Pro remain the picks; see the best AI image tools guide.
- The one to watch is the open-weight Dev release later in 2026. If a natively-multimodal model ships open, it changes the economics of AI video the way FLUX changed AI image.
The honest caveats
- Most of FLUX 3 isn’t broadly available. Only Video and Action are in early access (APIs/partners). Image is weeks out; open weights are months out. The “single model for image, video, and audio” is the architecture, not a product you can fully use today.
- The quality numbers are vendor-run. 77% over Runway and 93% over Luma come from Black Forest Labs’ own evaluations, with the lab controlling prompts and comparisons. Independent testing is the standard before treating them as fact.
- The comparison set is conspicuously incomplete. Veo — the model most often cited as the video-with-audio leader, and FLUX 3’s most direct rival on its own headline feature — wasn’t in the published head-to-heads. Draw the obvious inference cautiously, but draw it.
- Native audio is hard, and 20 seconds is short. Synchronized, convincing generated audio is genuinely difficult; early clips may impress in demos and disappoint on your specific prompts. And 20 seconds, while longer than many rivals, is still a short-form ceiling.
- Robotics claims are early partner tests. FLUX 3 Action being tested by Audi/mimic is a pilot, not a shipped capability. Treat it as a research direction, not a product.
The grounded summary: FLUX 3’s native-multimodal, jointly-trained approach — especially generating video and audio together — is a real architectural idea from a lab with a real track record, and the promised open-weight release could reshape AI video economics. But today it’s a partial early-access launch wrapped in vendor evals that skip the biggest competitors. Get on the list, keep your current tool, and judge it on your own footage when you can.
Frequently asked questions
What is FLUX 3?
FLUX 3 is a multimodal generative model from Black Forest Labs (the lab behind the FLUX image models), announced July 23, 2026. Its defining feature is that image, video, and audio are generated from a single set of weights trained jointly on all three, rather than separate models bolted together. FLUX 3 Video makes clips up to 20 seconds with synchronized native audio; a separate FLUX 3 Action model targets robotics.
What can FLUX 3 Video actually do?
It generates clips up to 20 seconds long with audio created alongside the visuals — dialogue, sound effects, and ambient noise — and supports text-to-video, image-to-video, and video-to-video, plus keyframe-controlled transitions and multilingual dialogue. The native-audio part is the notable bit: sound is generated jointly with the video rather than added afterward.
Is FLUX 3 better than Runway or Veo?
In Black Forest Labs' own evaluations, human reviewers preferred FLUX 3 over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93%. Those are vendor-run evals, not independent, and Google's Veo — the leading video-with-native-audio model, which is FLUX 3's own headline feature — wasn't part of the published comparison. Treat the numbers as a strong claim to verify, not a settled ranking — test on your own prompts once you have access.
Can I use FLUX 3 right now?
Only partly. At launch, FLUX 3 Video and FLUX 3 Action are in early access through APIs and select partners. Image generation is coming 'in the coming weeks,' and the open-weight 'Dev' version isn't due until later in 2026. So the headline 'one model for image, video, and audio' is the architecture; the shipping products today are Video and the robotics-focused Action model.
Why does 'natively multimodal' matter?
Most AI-media tools stitch separate image, video, and audio models behind one interface, which causes mismatches — audio that doesn't quite fit the video, or style drift between stills and motion. Training one model jointly on all three aims to produce output that's coherent across modalities by design, especially synchronized sound. If it works as claimed, it's a genuine architectural step, not just a bigger model.
Sources
- Black Forest Labs Unveils FLUX 3, A New Multimodal Frontier Model For Visual Intelligence (GlobeNewswire / Black Forest Labs)
- Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start (VentureBeat)
- Black Forest Labs' FLUX 3 Promises Images, Video and Audio; Which Features Actually Work Today (IBTimes)
- Black Forest Labs Unveils FLUX 3 AI: Ditches Stills for Video—And Robot Hands (Yahoo Tech)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.