FLUX 3 Image: Black Forest Labs Bets on Pixel-Perfect Layout
With FLUX 3 Image, you no longer just describe an image: you draw it box by box, then edit it without the rest moving.

Who covers it
In brief
Black Forest Labs introduces FLUX 3 Image, the image branch of its new multimodal FLUX 3 model, built around bounding boxes: each element is placed on a grid, then editable in isolation. The model renders in native 2K and 4K, accepts up to 10 references, and can be driven by a human or an agent. The bet: turn image generation into a controllable composition tool, not a slot machine.
🍺 Bar-stool version
Until now, generating an image with AI was a bit like ordering at a restaurant by describing the dish to a chef who doesn't speak your language: you get something, rarely what you wanted. Black Forest Labs now hands you a table plan: you draw boxes, say what's in each one, and the model arranges everything just right. And if you want to change the color of a surf suit, it won't repaint the beach in the process, which is already more polite than most software. For designers, agencies, and anyone churning out images at scale, that's the difference between a toy and a tool.
Key takeaways
- 1
FLUX 3 is Black Forest Labs' multimodal model (video, audio, images, actions); FLUX 3 Image is its dedicated brick for image generation and editing.
- 2
Composition relies on bounding boxes on a 0-to-1000 grid on both axes, regardless of aspect ratio, each box noted as [y_min, x_min, y_max, x_max].
- 3
A 'layout prompt' combines a global caption with a JSON element table (id, bbox, description), the caption referencing each element by its id.
- 4
Editing happens box by box, with multiple edits at once, promising that anything untouched stays exactly in place.
- 5
The model generates in native 2K and 4K, up to 5456 × 3072 pixels in the soba shop example, and combines up to 10 reference images.
- 6
Built for agents: an LLM can plan the layout on its own from a one-line prompt and a format, with each box still editable afterward.
- 7
Weights are offered under a commercial license to companies that want to fine-tune and deploy the model on their own infrastructure; public access goes through the Playground and the BFL API.
Composing rather than describing
The core of FLUX 3 Image is the bounding box. For each important element, you draw a box and describe its content, then a scene sentence ties everything together. The model renders the image while respecting each placement.
The grid is normalized from 0 to 1000 on both axes, regardless of the chosen ratio. The 'Festival of the Sun' example shows the logic: five rows in the table, a cream serif title typography, a coastal city, a concrete parabolic dome, silhouetted bathers, and a crowd on the beach.
The prompt format is explicit: a one-paragraph global caption, followed by a JSON table where each row carries an id, a box, and a description. The caption mentions each element by its identifier (animal_1, crowd_1…) wherever it appears.
Editing that breaks nothing
The second promise: retouch a finished image without distorting it. Each box can be redescribed, replaced, or moved, and several edits can be applied in a single pass.
Black Forest Labs emphasizes stability: whatever wasn't modified stays in place, allowing successive editing passes without the image drifting. Demos show a wetsuit and surfboard being recolored, or two divers added to an existing scene.
This is precisely the historical weak point of diffusion models, which tended to regenerate everything the moment you touched a detail.
The discreet role of the upsampler
Short queries go through a prompt upsampler, which turns them into dense captions, the format FLUX 3 was trained on. It can rewrite the caption around the boxes and suggest other elements.
But the rule is strict: every box drawn by the user reaches the model verbatim, with the same id and the same coordinates. Anything the upsampler adds survives only if the caption refers to it. A way to keep human input prioritized over automatic rewriting.
A model built for agents
Black Forest Labs presents FLUX 3 Image as 'designed for agents.' Just give it a one-liner and a format: an LLM plans the layout, writes the caption and the element table with semantic ids.
The user can then move boxes they're not happy with, and the model generates within the others exactly as the agent placed them. The 'Swan Lake seen from backstage' example illustrates this chain.
Without any boxes at all, the model remains a classic text-to-image system, with claimed solid prompt adherence and a native understanding of composition.
Intended uses and access
The highlighted use cases are those where many elements follow strict relationships: text composed around a photo, collages, thumbnail grids, editorial spreads, crowd scenes where every face has its place.
On the access side, you can draw boxes by hand in the Playground or send a layout prompt via the BFL API. Companies generating at scale can obtain the weights under a commercial license to fine-tune and host them in-house, upon request.
“Everything you didn't touch stays exactly where it was.”
“Every box you drew reaches the model verbatim, with the same id and the same coordinates.”
“Anything the upsampler adds survives only if the caption refers to it.”
Why it matters
FLUX 3 Image marks a shift in image generation: the battle is no longer only about render beauty, but about control. Bounding boxes, the JSON table, and localized editing turn the prompt into a specification, readable by a human as well as an agent, which fits perfectly with production workflows (advertising, press, e-commerce) where you want a precise, reproducible image. The structured format is also a natural fit for agents: an LLM is much better at writing JSON with coordinates than at guessing a model's aesthetics. That said, this remains a product page: no benchmarks, no pricing, no comparison with competitors, and the promise that 'everything else stays put' needs to be verified on hard cases. Also notable is the choice of a commercial weights license, on request, rather than broad openness: a signal that Black Forest Labs is now embracing an enterprise-vendor positioning. Finally, FLUX 3 is presented as multimodal (video, audio, actions), but this page only shows the image side: the rest of the family will be the real test of its ambition.
Free account
You just read an AI Sources article
Create a free account: a month of archives in full, your own sources summed up like this one, your notes and highlights.
For you
Put it to work on your sources.
Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.
For your team
The same machine, on your topics.
A space in your colours, your watch angles, your curators. Pilot open to three companies.
Read next
#superintelligenceYesterdayHinton, Bengio, OpenAI, and Anthropic Sign a Plan Against the Intelligence Explosion
Twenty-two researchers, including executives from OpenAI and Anthropic, argue that automating AI research could compress years of progress into mere months, and that governments aren't ready for it.
Source · Cambridge Programme on AI Science & Policy (CASP) · What if automating AI R&D triggers an intelligence explosion?
#apiYesterdayGPT-6: OpenAI's Playbook for Moving From Prototype to Production
Three models, five reasoning levels, two speeds and agents that work for hours: OpenAI publishes the manual for its new family.
Source · OpenAI · A model guide for the GPT-6 family
#open source2 OctClef: Cloudflare Launches Its Own Open Source Decision Models
With Clef and Clef-flash, Cloudflare enters the nascent market of "decision models" and adds a reinforcement learning fine-tuning platform to the mix.
Source · Cloudflare Blog · Introducing Clef: our open-source decision models, and new RL fine-tuning platform