AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Jul 24, 2026
·
videoimage-generationmodels

FLUX 3 generates image, video, and audio from one model — Black Forest Labs' bet on natively multimodal generation

TL;DR: Black Forest Labs — the lab behind the FLUX image models — unveiled FLUX 3 on July 23, calling it the first natively multimodal architecture: image, video, and audio from a single set of jointly-trained weights, not separate models stitched together. FLUX 3 Video makes clips up to 20 seconds with synchronized native audio (dialogue, SFX, ambient), plus text/image/video-to-video and keyframe transitions. In Black Forest Labs’ own evals, reviewers preferred it over Runway Gen-4.5 (77%) and Luma Ray 3.2 (93%) — vendor numbers, not independent, and notably Veo — the incumbent at video-with-native-audio, exactly FLUX 3’s headline feature — wasn’t in the comparison. What ships today: Video and a robotics-focused Action model in early access; image generation “in the coming weeks,” open weights later in 2026. What this means for you: a genuinely new architecture for AI video with native sound — promising, but verify the claims when you get access, because most of it isn’t broadly available yet.

What was announced

On July 23, 2026, Black Forest Labs unveiled FLUX 3, describing it as “the first natively multimodal architecture, bringing together image, video, audio and action generation within a single model.” The architectural claim is the story: rather than assembling separate image, video, and audio models behind a shared interface — the industry norm — FLUX 3 is trained jointly across all three modalities and generates them from one set of weights.

What that produces, per the company and VentureBeat:

What’s actually available: only Video and Action are in early access now (APIs and select partners). Image generation follows “in the coming weeks,” and the open-weight “Dev” version isn’t due until later in 2026. So the “one model for everything” headline is the architecture; the shipping products are Video and the robotics model.

Why this matters

1. Native multimodality is a real architectural shift, if it delivers. Today’s AI-media stacks glue an image model to a separate video model to a separate audio model. That’s why generated video so often has sound that doesn’t quite fit, or a look that drifts between the still and the motion. Training one model jointly on image, video, and audio targets exactly that seam — coherence across modalities by design, especially synchronized sound. Native audio in particular has been the missing piece in most AI video (Veo being the notable exception); a model that generates picture and sound together rather than dubbing after the fact is a genuine step, not a spec bump. The caveat is “if it delivers” — architecture claims are easy, coherent 20-second clips with convincing synced audio are hard.

2. The competitive framing targets Runway and Luma — but skips the biggest names. Black Forest Labs benchmarked against Runway Gen-4.5 and Luma Ray 3.2, and beat both handily in its own evals. Two honest asterisks: these are vendor-run preference tests (the lab chose the prompts, the comparisons, and what to publish), and Google’s Veo — the incumbent leader at video-with-native-audio, which is FLUX 3’s own headline feature — wasn’t in the comparison. That omission is the informative one: you benchmark against the rivals you can beat, and skipping the model that already does the exact thing you’re claiming as new is conspicuous. It doesn’t mean FLUX 3 is bad — it may be excellent — but “beats Runway and Luma” is a narrower claim than “best AI video model,” and the gap between those two statements is exactly where buyers get misled.

3. Black Forest Labs has credibility, which raises the stakes. This isn’t a random startup. BFL’s FLUX image models became a genuine open-weight standard for image generation, competing with Midjourney and Stable Diffusion. A lab with that track record moving into jointly-trained video+audio is worth taking seriously — and the promised open-weight Dev release later in 2026 matters most of all. If a high-quality, natively-multimodal model ships with open weights, it does to AI video what FLUX did to AI image and what DeepSeek did to text: pressure the closed leaders on price and access. That’s the development to watch, more than the launch-day evals.

4. The robotics angle is a tell about where “video models” are heading. FLUX 3 Action — a generation model pointed at robot manipulation rather than content — signals that “video generation” and “world modeling for robotics” are converging. A model that understands how objects move, sound, and interact well enough to make a convincing 20-second clip is learning something a robot can use. Most creators won’t care about Action directly, but it explains why a media lab is suddenly training on “action”: the same physical-world understanding powers both. It’s the clearest sign yet that generative video and embodied AI are becoming the same research problem.

5. “Limited release” is the pattern, and it shapes what you can trust. Like GPT-5.6’s staged rollout and Gemini 3.5 Pro’s endless preview, FLUX 3 launched as an announcement plus partial early access, not a product you can fully use. That’s now standard, and it means the marketing lands weeks before independent verification is possible. Read launch-day claims — including the 77%/93% numbers — as hypotheses until creators outside the partner program put it through real work.

What this means for you

The honest caveats

The grounded summary: FLUX 3’s native-multimodal, jointly-trained approach — especially generating video and audio together — is a real architectural idea from a lab with a real track record, and the promised open-weight release could reshape AI video economics. But today it’s a partial early-access launch wrapped in vendor evals that skip the biggest competitors. Get on the list, keep your current tool, and judge it on your own footage when you can.

Frequently asked questions

What is FLUX 3?

FLUX 3 is a multimodal generative model from Black Forest Labs (the lab behind the FLUX image models), announced July 23, 2026. Its defining feature is that image, video, and audio are generated from a single set of weights trained jointly on all three, rather than separate models bolted together. FLUX 3 Video makes clips up to 20 seconds with synchronized native audio; a separate FLUX 3 Action model targets robotics.

What can FLUX 3 Video actually do?

It generates clips up to 20 seconds long with audio created alongside the visuals — dialogue, sound effects, and ambient noise — and supports text-to-video, image-to-video, and video-to-video, plus keyframe-controlled transitions and multilingual dialogue. The native-audio part is the notable bit: sound is generated jointly with the video rather than added afterward.

Is FLUX 3 better than Runway or Veo?

In Black Forest Labs' own evaluations, human reviewers preferred FLUX 3 over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93%. Those are vendor-run evals, not independent, and Google's Veo — the leading video-with-native-audio model, which is FLUX 3's own headline feature — wasn't part of the published comparison. Treat the numbers as a strong claim to verify, not a settled ranking — test on your own prompts once you have access.

Can I use FLUX 3 right now?

Only partly. At launch, FLUX 3 Video and FLUX 3 Action are in early access through APIs and select partners. Image generation is coming 'in the coming weeks,' and the open-weight 'Dev' version isn't due until later in 2026. So the headline 'one model for image, video, and audio' is the architecture; the shipping products today are Video and the robotics-focused Action model.

Why does 'natively multimodal' matter?

Most AI-media tools stitch separate image, video, and audio models behind one interface, which causes mismatches — audio that doesn't quite fit the video, or style drift between stills and motion. Training one model jointly on all three aims to produce output that's coherent across modalities by design, especially synchronized sound. If it works as claimed, it's a genuine architectural step, not just a bigger model.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.