How to Write Bounding-Box Layout Prompts for FLUX 3 Image (Step by Step)

Why layout prompting beats plain prose

Text-to-image models infer composition from prose, which makes precise layouts hard to reproduce: "product on the left, headline top-right" comes out differently every run. FLUX 3 Image accepts a structured scene layout — a JSON element table plus a scene caption — so positions become explicit inputs instead of wishes. Each element gets an ID, a description, and a bounding box, and the model renders the final image around that plan.

The 0-to-1000 coordinate grid