Skip to content
The Executives BriefThe Executives BriefBeta

Black Forest Labs launches FLUX 3 with 20-second audio-video, but no pricing yet

Enter gated early access for video and action, while image waits, benchmarks stay preliminary, and costs are a mystery.

ByLama Al-RashidTechnology Correspondent, The Executives Brief
·5 min read
Black Forest Labs launches FLUX 3 with 20-second audio-video, but no pricing yet
Executive summary

Black Forest Labs (BFL) launched FLUX 3, a multimodal model that generates images or combined audio/video clips up to 20 seconds from a single prompt, plus an “Action” product line. For decision-makers, the limited release plus missing pricing, SLAs, and full benchmarks makes early procurement and vendor comparisons hard.

Black Forest Labs just launched FLUX 3, its multimodal step beyond image generation. The model can create combined audio and video clips up to 20 seconds from a single prompt, and BFL is also framing it as a shared backbone for “Action” and robotic vision, not separate models glued together behind one interface. The twist for enterprise buyers is that the company is not giving the usual procurement essentials right away. Pricing, production service-level commitments, evaluation methodology details, sample sizes, rater counts, and even image-model benchmarks are not public yet, and FLUX 3 is rolling out through gated early access for video and action.

Here is what you can access today, and what you cannot. FLUX 3 is offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and an upcoming open source FLUX 3 Dev. FLUX 3 Video (with optional native audio generation) and FLUX 3 Action enter a gated “Early Access” program now, and BFL says anyone can apply but BFL must approve. There is presently no public access through BFL’s API or those of partners yet. Meanwhile, FLUX 3 Image is expected to roll out in the coming weeks, followed by general availability. Also absent: downloadable weights and any open source license at launch. BFL says faster and open-weight versions will arrive later this year, and it describes FLUX 3 Dev as “open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction,” which is broader than previous FLUX Dev releases that covered images only.

Why this matters is not just “cool model, cool demos.” It changes how enterprises evaluate risk and cost. The release strategy mirrors what other frontier labs have done more recently in the U.S., including Anthropic and OpenAI, which also had limited rollouts that were tied to security concerns and government request. BFL is not announcing pricing or other operational guarantees at the same time it is asking customers to care about video generation with audio and action prediction. That means finance and procurement cannot yet answer the basic questions: what is the total cost of ownership, how do service levels behave in real production, and how do results compare when you actually run your own prompts and assets?

BFL does publish benchmark comparisons, but it keeps a big qualifier attached. The published results are described as a “preliminary evaluation of an early FLUX 3 candidate,” with full benchmark results and methodology to be published later during broader general availability. In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, BFL says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%. Those are real signals. But they are not yet the same thing as operational proof. The model “now entering early access” could differ from the preliminary checkpoint, and the published materials do not yet let independent teams reproduce the video comparisons.

The competitive context also gets messy fast, because access differs by region and by product packaging. Gemini Omni Flash is the closest large-platform analog to what FLUX 3 is attempting: multimodal input, video and audio-aware creation, and conversational editing. In BFL’s own measurement, Omni is indistinguishable from FLUX 3 on 10-second text-to-video quality. The practical difference is availability and editing permissions. Omni Flash is generally available via Google’s Gemini API at $0.10 per second of generated 720p video, which works out to around $1.00 for a 10-second clip. But for a German company’s home-market enterprise buyers in the European Economic Area, Switzerland, and the United Kingdom, editing uploaded video is unavailable to Omni Flash users. Editing video that the model itself generated is permitted. For European enterprises, that specific limitation can break workflows that depend on generative passes over existing footage.

Meanwhile, Seedance 2.0 is a reminder that regulation and enforcement shape the marketplace as much as model quality. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount, and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a “frozen product” to a preference statistic is not a useful enterprise decision by itself. It mostly tells you who is accessible, not who is most likely to win your internal review. In other words: FLUX 3’s early numbers may be impressive, but accessibility, licensing, and operating constraints often decide the shortlist.

Under the hood, BFL says it is not assembling separate systems. FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture that the company publicized in March 2026. BFL says it scaled up compute and data to train across video, images, and audio simultaneously, and testing showed that video generation and action prediction do not require separate foundations. Its pitch is simple: vision is the most signal-rich medium of the physical world, but vision alone is not the complete picture. BFL’s co-founder and CEO Robin Rombach said in the pre-release statement provided to VentureBeat that “Joint training within one unified architecture is what will get us there,” because each training modality strengthens the others. He added that audio conveys timing and events that elude vision, and language conveys goals and abstractions pixels cannot express as easily. The company’s framing is that “True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results.”

So what should peers in the same room do with this? The strategy is clear: BFL wants enterprises to treat creative generation, simulation, computer use, and robotics as connected applications of one “visual intelligence” capability. But the near-term buying problem is equally clear: limited release, no published SLA or pricing, no downloadable weights, and only preliminary benchmarks make it harder to underwrite production adoption today. The market is moving toward multimodal, real-world-aware models. FLUX 3 is positioned to be part of that story, but BFL is asking customers to commit before it reveals the numbers that procurement needs most.

Executive ActionsLocked

This story's Key Insights and Take-aways are locked.

Create a free account to unlock Executive Actions for one credit.

Register to Unlock

Always free for Executives Club members. Join the Club

More in Technology