FLUX 3 Image

FLUX 3 Image is Black Forest Labs' image generation and editing model built on the multimodal FLUX 3 backbone shared with FLUX 3 Video. It combines text-to-image synthesis, precise local editing, multi-reference composition, bounding-box placement, and native 4K output. The model preserves identities and fine details across references, supports targeted changes and in-place text editing, and renders accurate typography in text-heavy layouts across a broad range of visual styles.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
Bounding boxes with FLUX 3 Image How to place and edit elements by region with FLUX 3 Image: the 0 to 1000 grid, composing a layout from scratch, moving a marked element, and removing one.
Editing images with FLUX 3 Image How to edit an image with FLUX 3 Image from a single reference: scoping the instruction, swapping materials and light, removing objects, reframing, and chaining edits.
Multi-reference composition with FLUX 3 Image How to combine up to ten reference images with FLUX 3 Image: assigning a role to each one, holding a face or a product across scenes, and borrowing a look.
Prompting FLUX 3 Image How to prompt FLUX 3 Image for text-to-image: the five-part prompt structure, the detail that prompt expansion cannot invent for you, camera direction, and output size.
Rendering text with FLUX 3 Image How to render readable text with FLUX 3 Image: quoting the exact string, directing the typography, building a text hierarchy, and editing copy inside an existing image.