live
MODEL IDbfl:flux@3-image

FLUX 3 Image

Black Forest Labs
by
Available with Zero Data Retention

FLUX 3 Image is Black Forest Labs' image generation and editing model built on the multimodal FLUX 3 backbone shared with FLUX 3 Video. It combines text-to-image synthesis, precise local editing, multi-reference composition, bounding-box placement, and native 4K output. The model preserves identities and fine details across references, supports targeted changes and in-place text editing, and renders accurate typography in text-heavy layouts across a broad range of visual styles.

FLUX 3 Image

Multi-reference composition with FLUX 3 Image

How to combine up to ten reference images with FLUX 3 Image: assigning a role to each one, holding a face or a product across scenes, and borrowing a look.

Introduction

inputs.referenceImages takes up to ten images in one request, and the model reads them as raw material rather than as a stack of layers. Nothing about the array says which image is the subject, which is the backdrop and which is only there for its color treatment. The prompt assigns those roles, by index.

That is the whole technique. Once each image has a job in the sentence, a face, a garment and a location that were never photographed together come back as one frame.

Built from 3 references
  • The face
  • The jacket
  • The location

The headshot was lit flat in a studio and the location is overcast, so the prompt tells the model which of those two lighting conditions wins. That instruction is what stops the composite from reading as a cutout.

This guide covers how to write the role assignment, how to hold a face or a product steady across a set, how to use one reference purely for its look, and what the references are worth feeding in.

Assigning a role to each image

Reference images are numbered in the order you send them, and the prompt refers to them as image 1, image 2 and so on. The references are used either way: that is what sending them is for. What the roles decide is where the sofa lands and which lighting wins, and that is what makes the same request repeatable across a catalog.

The prompt produces the same catalog shot on every run, because four separate instructions are doing four separate jobs:

  • The setting: "use image 2 as the room"
  • The subject and its placement: "place the sofa from image 1 against the wall opposite the window, sitting flat on the herringbone floor"
  • What must survive the transfer: "keep the sofa exact including its mustard velvet, its rounded arms and its tapered walnut legs"
  • Which lighting wins: "match the soft daylight of image 2, with a soft contact shadow under the legs"

That last pair is what separates a composite from a paste. A product carries the light of the studio it was shot in, and unless the prompt says to relight it, some of that studio comes along.

Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'bfl:flux@3-image',
  positivePrompt: 'Use image 2 as the room. Place the sofa from image 1 against the wall opposite the window, squared up to the camera and sitting flat on the herringbone floor. Keep the sofa exact including its mustard velvet, its rounded arms, its single bench cushion and its tapered walnut legs. Match the soft daylight of image 2, falling on the sofa from the left, with a soft contact shadow under the legs. Wide shot at chest height, 24mm lens. Photoreal interior photography.',
  inputs: {
    referenceImages: [
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/e5f6a7b8-c9d0-1234-ef12-345678901234.jpg',
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/f6a7b8c9-d0e1-2345-f123-456789012345.jpg'
    ]
  },
  width: 1248,
  height: 832
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "bfl:flux@3-image",
            "positivePrompt": "Use image 2 as the room. Place the sofa from image 1 against the wall opposite the window, squared up to the camera and sitting flat on the herringbone floor. Keep the sofa exact including its mustard velvet, its rounded arms, its single bench cushion and its tapered walnut legs. Match the soft daylight of image 2, falling on the sofa from the left, with a soft contact shadow under the legs. Wide shot at chest height, 24mm lens. Photoreal interior photography.",
            "inputs": {
                "referenceImages": [
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/e5f6a7b8-c9d0-1234-ef12-345678901234.jpg",
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/f6a7b8c9-d0e1-2345-f123-456789012345.jpg"
                ]
            },
            "width": 1248,
            "height": 832
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "imageInference",
      "taskUUID": "c3d4e5f6-a7b8-9012-cdef-123456789012",
      "model": "bfl:flux@3-image",
      "positivePrompt": "Use image 2 as the room. Place the sofa from image 1 against the wall opposite the window, squared up to the camera and sitting flat on the herringbone floor. Keep the sofa exact including its mustard velvet, its rounded arms, its single bench cushion and its tapered walnut legs. Match the soft daylight of image 2, falling on the sofa from the left, with a soft contact shadow under the legs. Wide shot at chest height, 24mm lens. Photoreal interior photography.",
      "inputs": {
        "referenceImages": [
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/e5f6a7b8-c9d0-1234-ef12-345678901234.jpg",
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/f6a7b8c9-d0e1-2345-f123-456789012345.jpg"
        ]
      },
      "width": 1248,
      "height": 832
    }
  ]'
runware run bfl:flux@3-image \
  positivePrompt="Use image 2 as the room. Place the sofa from image 1 against the wall opposite the window, squared up to the camera and sitting flat on the herringbone floor. Keep the sofa exact including its mustard velvet, its rounded arms, its single bench cushion and its tapered walnut legs. Match the soft daylight of image 2, falling on the sofa from the left, with a soft contact shadow under the legs. Wide shot at chest height, 24mm lens. Photoreal interior photography." \
  inputs.referenceImages.0=https://im.runware.ai/image/os/a14d18/ws/2/ii/e5f6a7b8-c9d0-1234-ef12-345678901234.jpg \
  inputs.referenceImages.1=https://im.runware.ai/image/os/a14d18/ws/2/ii/f6a7b8c9-d0e1-2345-f123-456789012345.jpg \
  width=1248 \
  height=832
{
  "taskType": "imageInference",
  "taskUUID": "c3d4e5f6-a7b8-9012-cdef-123456789012",
  "model": "bfl:flux@3-image",
  "positivePrompt": "Use image 2 as the room. Place the sofa from image 1 against the wall opposite the window, squared up to the camera and sitting flat on the herringbone floor. Keep the sofa exact including its mustard velvet, its rounded arms, its single bench cushion and its tapered walnut legs. Match the soft daylight of image 2, falling on the sofa from the left, with a soft contact shadow under the legs. Wide shot at chest height, 24mm lens. Photoreal interior photography.",
  "inputs": {
    "referenceImages": [
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/e5f6a7b8-c9d0-1234-ef12-345678901234.jpg",
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/f6a7b8c9-d0e1-2345-f123-456789012345.jpg"
    ]
  },
  "width": 1248,
  "height": 832
}
Response
{
  "data": [
    {
      "taskType": "imageInference",
      "taskUUID": "c3d4e5f6-a7b8-9012-cdef-123456789012",
      "imageUUID": "a7b8c9d0-e1f2-3456-1234-567890123456",
      "imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/a7b8c9d0-e1f2-3456-1234-567890123456.jpg"
    }
  ]
}

Holding a product across a set

One reference is enough when the job is consistency rather than combination. Send the packshot, then change everything around it. The clause that does the work is the list of features that have to survive, and it should name the things a buyer would notice if they moved: the color, the closure, the proportions, the label.

Three scenes and one product. Each prompt relights the bottle for its scene while the sage green, the black cap and the proportions stay put, which is what lets the set run as one campaign. The same pattern with a face instead of a bottle is how a spokesperson holds across a series.

A generated product is a likeness, not a duplicate. Fine print, a logo lockup and an exact color reference will not survive a transfer reliably, so anything that has to match a real SKU down to the label belongs in a composite you control rather than in the generation.

Borrowing a look

Style is just another role. Send the picture you want to keep and the picture whose treatment you want, then say plainly which is which and that the layout must not move.

Built from 2 references
  • The scene
  • The treatment

Two references arrive and only one of them is content, so the prompt has to say which. Naming the layout as something to preserve is what keeps the treatment reference from contributing a subject, which is why the prompt ends on "do not change the layout or add anything to the scene" rather than on the description of the grain.

What the references are worth

A reference has to be at least 128 × 128 pixels, and that is the only size that gets a request rejected. Anything larger is accepted and scaled down to fit 6256 pixels on the long side and 16 MP in total, so a reference past that ceiling costs upload time on every request in the batch and buys nothing.

What does pay off is what the reference contains. A packshot on a clean background transfers better than the same product buried in a busy scene, because the model has less to disentangle before it can carry the object across. The same holds for a face: an evenly lit frontal portrait carries further than a three-quarter shot in mixed light.

Reference images are not billed. A request with one reference and a request with ten cost the same as the text-to-image call at that resolution tier, which makes multi-reference work cheap to iterate on.

When a request carries references, resolution becomes available as an alternative to width and height: it sets the output tier while the references supply the aspect ratio. Send one or the other, never both. A request that sends neither is rendered at the 1K tier.

Tips for best results

  1. Refer to images by index. "Image 1", "image 2". The array order is the only handle you have, and a prompt that says "the jacket" leaves the model matching by description instead.

  2. Give every reference a job. An image in the array with no role in the sentence still influences the output, usually by leaking color or texture into the result.

  3. Name what has to survive. List the features a buyer would notice: the color, the closure, the proportions, the trim. General instructions to stay faithful do less than four concrete nouns.

  4. Say which lighting wins. A product shot in a studio and dropped into daylight will keep its studio light unless the prompt hands the scene's lighting to it.

  5. Ask for a contact shadow. Where the object meets the floor or the surface is where a composite gives itself away.

  6. Shoot references clean. A plain background and even light on the reference transfer further than a busy frame, whatever the resolution.

  7. Keep references under 16 MP. They are downscaled past that, so the extra pixels only cost upload time. Below 128 × 128 the request is rejected.