FLUX 3: What Is New, How to Access It, and What It Costs

FLUX 3 is Black Forest Labs' attempt at one model across images, video, audio and robot action prediction. Here is what the official page claims and how to try it.

The short readThe interesting claim here is consolidation: one model covering images, video with native audio, and action prediction for robotics, rather than four models stitched together behind an interface. If you generate images or short clips, that matters mainly for consistency between them. If you were hoping for a small local model, this is not that conversation, and the official page does not commit to weights for this release. Pricing is not on the landing page, so check the official pricing page before planning a budget. Synexa is the route we use when the priority is calling FLUX-family models from code with one endpoint and paying per run instead of managing another vendor account.

Try Synexa → Official site

One model instead of four

The framing on the official page is deliberate and worth taking seriously: a single multimodal model spanning image, video, audio and action prediction, described as producing results truer to life across styles. Most stacks today reach that coverage by routing between specialists, an image model here, a video model there, a separate speech system for the soundtrack. Consolidation buys coherence. When the same model draws the frame and generates the sound that belongs in it, the pieces tend to agree with each other, which is the part that breaks first when you assemble a clip from three unrelated systems. Whether the unified approach beats a strong specialist on any single axis is an open question the marketing page does not settle, and only your own prompts will.

Video with sound in one pass

The video description is the most specific part of the page. Clips run up to twenty seconds in a single generation, which is longer than the few-second outputs people have become used to, and the site says you can start from text, from an image, or from keyframes and get multiple shots in one take. Audio is optional and generated alongside the frames, covering multilingual speech, effects and ambience rather than music alone. Keyframe conditioning is the practical detail for anyone doing production work, because it is the difference between accepting whatever the model imagines and steering a shot toward something you have already designed. A separate video upscaler to 2K and 4K is advertised on the same site as its own tool.

Images and the text rendering claim

For still images the page emphasises grounding in the real world, a wide range of styles, complex prompt handling, and highly accurate text rendering. That last item is the one to test first if it matters to you, because rendered text is where image models still fail visibly and where the gap between a demo and your use case is widest. Set up a fair comparison rather than a flattering one: a poster with a real brand name, a UI mockup with several labels, a product shot with small print on the packaging. A model that survives all three is genuinely useful for marketing work. One that only manages a single large word is fine for hero images and will frustrate you everywhere else.

Action prediction is a different audience

The fourth modality is not for creative teams at all. The site describes it as unifying perception, simulation and execution for robotics, taking visual observations plus text instructions and predicting the desired physical outcome both visually and as robot control actions, with a dedicated blog post behind it. This is an unusual thing to ship next to a text-to-image product, and it explains the emphasis on being grounded in the real world: a model that has to predict what physically happens next is under a different kind of pressure than one that only has to look plausible. For anyone evaluating the creative side, it is context rather than a feature, and the research pages are the place to read further.

Access, pricing and calling it from code

Getting in starts at the vendor dashboard, with a playground for trying prompts, an API section and a pricing page in the navigation, plus an enterprise track and a contact route for sales. No figures appear on the landing page itself, so treat any price you read on a third-party site as unverified and read the official pricing page for current rates. Note also that the site keeps an open weights section, but the landing page does not state what is included there for this release, so do not assume local deployment is possible. Synexa is the alternative path when you want FLUX-family models, plus video and audio models, behind one REST endpoint and a Python SDK with per-run billing.

The specifics the page commits to

Clips up to twenty seconds

A single generation is described as producing up to twenty seconds, with multiple shots in one take, which changes what you can do without stitching outputs together afterwards.

Audio generated with the frames

Optional sound covering multilingual speech, effects and ambience, produced alongside the video rather than added afterwards, which is the point where most hand-assembled pipelines stop sounding like they belong together.

Text rendering in images

Highly accurate text rendering is claimed, alongside broad style coverage and complex prompt handling. Worth testing early with real labels and small print rather than a single large word.

Robotics action prediction

A modality aimed at robot control: visual observations and text instructions in, predicted physical outcomes and control actions out. Documented in a separate post on the vendor's blog.

Going direct, or going through Synexa

What you needFLUX 3 from the vendorSynexa
Trying prompts by handPlayground in the vendor dashboardNot the focus; built for code
Calling it from an applicationVendor API, per the site navigationOne REST endpoint plus a Python SDK
Mixing image, video and audio modelsWithin this model familyMultiple model families behind one API
Which FLUX versions are availableListed on the official models pageCheck the Synexa model list for what is live
Billing modelSee the official pricing pagePay per run
Running the weights yourselfNot stated for this release on the landing pageNot offered; it is a hosted API
Enterprise routeContact sales link on the siteStandard API access

Getting from curious to a real test

  1. Start in the playground
    Use the vendor dashboard to try prompts by hand before writing any code. Ten minutes there tells you whether the output style suits your work, which no benchmark will.
  2. Test the hard cases early
    Run rendered text, a twenty second clip and a keyframe-driven shot on day one. Those are the claims with the most to prove, and failures there reshape your plans.
  3. Read the pricing page properly
    Find current rates on the official page and work out the cost of a realistic week, not a single call. Video and audio generation prices differ sharply from still images.
  4. Decide how you will call it
    If several model families will end up in your product, compare a direct vendor account against Synexa's single endpoint with per-run billing before you write the integration twice.

Questions about FLUX 3

What is new in FLUX 3?

The headline change is scope: one model covering images, video with optional generated audio, and action prediction for robotics, presented as a single system rather than separate products. The video side is described as up to twenty seconds per generation with multiple shots, started from text, an image or keyframes.

How do I get access to FLUX 3?

Through the vendor's dashboard, which offers a playground for manual prompting and an API for programmatic use, with an enterprise track and a sales contact for larger deployments. Start in the playground to judge output quality before committing engineering time to an integration.

How much does FLUX 3 cost?

No pricing appears on the landing page, so any figure quoted elsewhere should be treated as unverified. The official pricing page carries current rates. Budget by modality rather than by a single number, since video with audio and still image generation sit at very different price points.

Are the weights available to download?

The site keeps an open weights section in its navigation, but the landing page does not state what it includes for this release, so do not assume local deployment is possible. Check the official models and open weights pages directly before designing anything around self-hosting.

Can it generate video with sound?

That is how the site describes it: audio is optional and produced alongside the frames, covering multilingual speech, sound effects and ambience. Generating both together is the interesting part, because separately produced audio is where hand-assembled pipelines usually stop feeling coherent.

What is the easiest way to call it from code?

Either the vendor API directly, or a hosted layer if you expect to use several model families. Synexa exposes FLUX-family models along with video and audio models through one REST endpoint and a Python SDK, billed per run, which saves maintaining a separate account and client for each vendor.

One endpoint for image, video and audio models

Synexa runs FLUX-family, video and audio models behind a single REST API and a Python SDK, billed per run. Wire up one client instead of a new account for every vendor you want to try.

Try Synexa →