FLUX 3 is now inside Luma, and it brings native audio to AI video for the first time
Black Forest Labs' first general multimodal model lands inside Luma Dream Machine, delivering video with built-in audio and a pipeline that spans image, video, sound, and action prediction in one place.
Something significant happened quietly on 12 August. Black Forest Labs, the team behind the FLUX image models that quietly became the default for serious AI image work, dropped their first model that spans every modality at once: image, video, audio, and action prediction. It landed inside Luma the same day, available immediately. That combination, a new model architecture from one of the most trusted names in generative image work, arriving pre-integrated into one of the busiest AI video platforms, is worth sitting with for a moment.
The model is called FLUX 3, and it is a significant departure from what Black Forest Labs has shipped before. Every FLUX release to date has been an image model. FLUX 3 is the studio's first general model, meaning it is trained to reason across modalities natively, not assembled by bolting separate components together. The early delivery inside Luma is video with native audio. That phrase matters: native audio means the audio is not a post-generation add-on synchronised after the fact, it is generated as part of the same pass. For anyone who has spent time trying to marry AI video to separately generated sound, the friction that workflow creates is immediately obvious.
What Luma is actually shipping today
@LumaLabsAI announced FLUX 3 as live in Luma with video and native audio available today. The announcement notes that more is coming soon, which is the standard holding phrase for the image, action-prediction, and broader multimodal capabilities the model is described as supporting. So the current delivery is video-with-audio, not the full suite.
It arrives alongside Luma Scenes, also announced this week, which is the platform's new scene-by-scene refinement workflow built on the Uni-1 model. FLUX 3 appears to sit as a separate generation option within Luma's toolkit, though the exact interface placement and how it relates to Uni-1 in practice has not been spelled out.
What is still unknown
Quite a lot, as is common with day-one launches. No pricing has been announced for FLUX 3 generations specifically, and it is not clear whether it will sit within existing Luma credit tiers or carry separate costs. No maximum clip length has been published. No benchmark comparisons to existing Luma models or to competitors have been released, so the quality claims are marketing-adjacent until independent tests emerge.
The action-prediction capability is intriguing, particularly for anyone interested in controllable motion, but there is no date for when that becomes available. The image generation mode similarly has no release window beyond "more coming soon." Whether FLUX 3 in Luma will offer the same LoRA fine-tuning flexibility that made FLUX image models so attractive to the creative community is also unaddressed.
Why this matters for working filmmakers
The practical gap that native audio generation addresses is real. Current AI video workflows almost universally treat audio as a second step: generate the clip, export, load into a separate tool, generate or sync sound, reassemble. That process is tedious and introduces alignment problems, particularly for anything with speech, where lip sync from a separately generated audio track rarely holds across a full clip.
If FLUX 3 genuinely produces coherent audio as part of the same generation, it compresses multiple production steps into one. The caveat is that first-generation multimodal audio from any system tends to be serviceable rather than broadcast-ready, and nothing in the announcement speaks to audio quality, voice consistency, or music generation capability. Those details will only emerge from filmmaker testing.
The Black Forest Labs pedigree matters here. The FLUX image models, particularly FLUX.1, became the go-to for photorealistic and stylised image work partly because the team had genuine research depth and shipped incrementally rather than overclaiming. That track record earns some credibility for FLUX 3, but it also means the community will hold the audio quality to a high standard quickly.
For filmmakers currently inside the Luma ecosystem, the obvious first test is a short dialogue scene: generate a clip with a speaking character and assess whether the audio tracks naturally, whether the voice quality holds, and whether any ambient sound coherently matches the visual environment. That will tell you more than any benchmark.
For filmmakers not yet on Luma, FLUX 3 is now a reason to open an account. One-stop video-with-audio generation from a credible model architecture is exactly the kind of workflow compression that makes a platform worth evaluating, even before the full multimodal suite arrives.
---
Sources: FLUX 3 live in Luma announcement by @LumaLabsAI; Luma Scenes announcement by @LumaLabsAI