Black Forest Labs' FLUX 3 Image adds bounding-box layout control to text-to-image
- FLUX 3 Image generates an image from bounding boxes plus a scene prompt: the sample "Le Festival du Soleil" was built from five boxes and an element table covering the title text, a coastal town, a pale concrete dome, swimmers in dark water and a crowd seated on the beach.
- Each box is a JSON entry with an id, bbox coordinates and a description, for example dome_1 at [250, 150, 650, 850] described as "a massive, smooth parabolic dome of pale concrete", so placement is declared before generation runs.
- Text is placed the same way: box Fr_Text_1 at [10, 200, 170, 800] puts "LE FESTIVAL DU SOLEIL" in a thin serif typeface in light cream, with a Replay control and an element table on the demo.
- The page describes the model as text-to-image with strong prompt following and a native understanding of composition, with samples including a woman in a striped mohair hood, a fashion collage titled XU ZHI, four prototype race cars on a wet track and an exploded-view pencil drawing of a wooden chair.
Hacker News opinions
Looks cool, and I'm eternally grateful these models are marked for open weight releases. That's the part I actually care about.
The steering you get from the latest Gemini releases has been nice to work with, and it's good to see declarative controls built into the API here, with an open model coming soon.
The UX looks amazing and very steerable, congrats to the team for focusing on the interface. Chats are awful user interfaces.