FLUX 3 Image

FLUX 3 Image is Black Forest Labs' image generation and editing model built on the multimodal FLUX 3 backbone shared with FLUX 3 Video. It combines text-to-image synthesis, precise local editing, multi-reference composition, bounding-box placement, and native 4K output. The model preserves identities and fine details across references, supports targeted changes and in-place text editing, and renders accurate typography in text-heavy layouts across a broad range of visual styles.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
← All GuidesPrompting
How to prompt FLUX 3 Image for text-to-image: the five-part prompt structure, the detail that prompt expansion cannot invent for you, camera direction, and output size.
Introduction
FLUX 3 Image takes a prompt and an output size and returns one image. There is no negative prompt, no seed, and no step count, so the prompt and the dimensions are the entire control surface. Everything you want in the frame has to be said, and everything you want out of it has to be said as a positive statement about what is there instead.
The model also expands your prompt before it generates, and there is no switch to turn that off. A short prompt gets filled in with choices you didn't make. A detailed one leaves less room for the model to invent.

A woman in her thirties holding a forearm plank on a charcoal mat in a bright studio, weight on her elbows, hips level, eyes down. She wears a slate blue cropped top and matching leggings. Pale birch floor, a white brick wall behind her, a rack of dumbbells out of focus at the far end. Broad daylight from a window on the left, a soft shadow along the mat. Low side-on shot at floor height, 35mm lens, the subject filling the right two thirds of the frame. Photoreal fitness photography, cool neutral palette.
This guide covers the prompt structure the model reads best, how to direct the camera, how to pick an output size, and how to iterate when there is no seed to hold onto.
The five parts of a prompt
Write the prompt in the order the model reads it: subject, action, location, composition, style. Each part answers a question the model would otherwise answer for you, and putting them in that order means the scene is established before the camera is placed on it.

A man in his forties pouring coffee at a kitchen island, both hands on the pot, looking down at the cup. An open-plan apartment with pale oak cabinets, a white quartz worktop and a sliding glass door onto a balcony. Morning sun from the right throwing a long rectangle of light across the floor. Wide shot at chest height, 24mm lens, the island running diagonally from the lower left. Photoreal interior photography, warm neutral palette.
The action clause is the one most people drop. "A man at a kitchen island" and "a man pouring coffee at a kitchen island, both hands on the pot" produce very different postures, and the second is the one you can brief a shoot around.
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'bfl:flux@3-image',
positivePrompt: 'A man in his forties pouring coffee at a kitchen island, both hands on the pot, looking down at the cup. An open-plan apartment with pale oak cabinets, a white quartz worktop and a sliding glass door onto a balcony. Morning sun from the right throwing a long rectangle of light across the floor. Wide shot at chest height, 24mm lens, the island running diagonally from the lower left. Photoreal interior photography, warm neutral palette.',
width: 1248,
height: 832
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "bfl:flux@3-image",
"positivePrompt": "A man in his forties pouring coffee at a kitchen island, both hands on the pot, looking down at the cup. An open-plan apartment with pale oak cabinets, a white quartz worktop and a sliding glass door onto a balcony. Morning sun from the right throwing a long rectangle of light across the floor. Wide shot at chest height, 24mm lens, the island running diagonally from the lower left. Photoreal interior photography, warm neutral palette.",
"width": 1248,
"height": 832
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "imageInference",
"taskUUID": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"model": "bfl:flux@3-image",
"positivePrompt": "A man in his forties pouring coffee at a kitchen island, both hands on the pot, looking down at the cup. An open-plan apartment with pale oak cabinets, a white quartz worktop and a sliding glass door onto a balcony. Morning sun from the right throwing a long rectangle of light across the floor. Wide shot at chest height, 24mm lens, the island running diagonally from the lower left. Photoreal interior photography, warm neutral palette.",
"width": 1248,
"height": 832
}
]'runware run bfl:flux@3-image \
positivePrompt="A man in his forties pouring coffee at a kitchen island, both hands on the pot, looking down at the cup. An open-plan apartment with pale oak cabinets, a white quartz worktop and a sliding glass door onto a balcony. Morning sun from the right throwing a long rectangle of light across the floor. Wide shot at chest height, 24mm lens, the island running diagonally from the lower left. Photoreal interior photography, warm neutral palette." \
width=1248 \
height=832{
"taskType": "imageInference",
"taskUUID": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"model": "bfl:flux@3-image",
"positivePrompt": "A man in his forties pouring coffee at a kitchen island, both hands on the pot, looking down at the cup. An open-plan apartment with pale oak cabinets, a white quartz worktop and a sliding glass door onto a balcony. Morning sun from the right throwing a long rectangle of light across the floor. Wide shot at chest height, 24mm lens, the island running diagonally from the lower left. Photoreal interior photography, warm neutral palette.",
"width": 1248,
"height": 832
}Response
{
"data": [
{
"taskType": "imageInference",
"taskUUID": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"imageUUID": "f1e2d3c4-b5a6-7890-1234-567890abcdef",
"imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/f1e2d3c4-b5a6-7890-1234-567890abcdef.jpg"
}
]
}Detail beats adjectives
Prompt expansion fills every gap you leave, so the gaps are where the output drifts. "Nice lighting" becomes whatever the model considers nice that run. "Hard afternoon light from the upper left casting a short sharp shadow to the right" becomes a lighting setup you can reproduce in the next shot.

A running shoe on a surface, nice lighting.

A single road running shoe in white mesh with a coral midsole and a black outsole, laces tied, standing on a pale concrete slab. A plain bone-white wall behind it. Hard afternoon light from the upper left casting a short sharp shadow to the right. Three-quarter view at shoe height, 50mm lens, the shoe centered with clear space above it. Photoreal footwear e-commerce photography, clean high-key palette.
The second prompt names the colorway, the surface, the direction of the light, the length of the shadow, the camera height and the lens. None of those are adjectives. Every clause is a decision taken away from the expander, which is what makes the shot repeatable across a catalog.
The useful ones to nail down, in rough order of how much they change the frame:
- Materials and finish: "white mesh", "brushed stainless", "heavy cable knit"
- Light direction and hardness: "hard light from the upper left", "flat overcast daylight", "a single warm lamp on the right"
- Camera height and lens: "at shoe height, 50mm", "wide at standing height, 24mm"
- Where the subject sits in the frame: "centered with clear space above", "filling the right two thirds"
Directing the camera
Photographic vocabulary is the fastest lever in the prompt. The model responds to angle, lens length and distance as instructions rather than as flavor, so one subject can carry a whole set of shots.

A frosted glass serum bottle with a matte white dropper cap standing on a pale travertine block, a plain sand-colored wall behind it. Hard morning light from the right throwing a long shadow to the left. Low angle from just below the block, 50mm lens, the bottle rising against the wall. Photoreal beauty product photography, warm neutral palette.

A frosted glass serum bottle with a matte white dropper cap lying on a pale travertine slab, the dropper resting beside it, a sprig of eucalyptus in the lower left. Soft diffused light from above. Flat overhead shot looking straight down, 50mm lens, the bottle running diagonally across the frame. Photoreal beauty product photography, warm neutral palette.

The neck and dropper cap of a frosted glass serum bottle filling the frame, a single bead of serum hanging from the glass pipette, the label edge just visible below. Soft directional light from the left picking out the frosted texture. Macro shot at 100mm, extremely shallow depth of field, the background falling away to a warm blur. Photoreal beauty product photography, warm neutral palette.
The three prompts share a subject and a palette and differ only in the composition clause. That is the pattern for building a product detail page from one brief: fix the subject and the style, then vary the camera.
Choosing the output size
width and height accept a fixed set of pairs rather than a free range. The pairs are organized as four size tiers crossed with fifteen aspect ratios, from 0.75K up to native 4K, and every pair is delivered exactly as requested. The full table is in the model's API reference.
The tier decides two things. It sets how much fine detail survives, and it is what the request is billed on, with a steep step up at the top of the range. Reach for 4K when the image will be cropped, zoomed, or printed, and stay at 1K for layout comps and drafts.

A woman in her twenties seated on a pale linen bench, turned three-quarters to camera, wearing a heavy cream cable-knit sweater with a rolled neck and a thin gold chain. Her hands rest in her lap. A plain putty-gray wall behind her. Soft window light from the left, a gentle falloff across the knit. Chest-up shot at eye height, 85mm lens, the sweater filling the lower half of the frame. Photoreal editorial fashion photography, muted natural palette, fine fabric texture.

A woman in her twenties seated on a pale linen bench, turned three-quarters to camera, wearing a heavy cream cable-knit sweater with a rolled neck and a thin gold chain. Her hands rest in her lap. A plain putty-gray wall behind her. Soft window light from the left, a gentle falloff across the knit. Chest-up shot at eye height, 85mm lens, the sweater filling the lower half of the frame. Photoreal editorial fashion photography, muted natural palette, fine fabric texture.
The crop is one fifth of the frame at full pixel size. The cable stitches hold their twist and fiber texture at that magnification, which is the practical argument for generating at 4K: a single render covers the hero, the detail crop and the print asset instead of three separate jobs.
resolution sets the tier on its own, but it only applies when the request carries reference images, where the references supply the aspect ratio. A text-to-image request sizes itself with width and height, and falls back to 1024 × 1024 when neither is sent. See the multi-reference composition guide.
Iterating without a seed
The same prompt sent twice returns two different images, and there is no seed to pin one down. The response carries the seed that was used, but it cannot be sent back.

A tan leather weekender bag standing upright on a bleached oak floor beside a white wall, its brass zip catching the light, twin handles folded over the top. Late afternoon light from a window on the left, a soft shadow to the right. Straight-on shot at bag height, 50mm lens, the bag centered with clear floor below. Photoreal accessories e-commerce photography, warm neutral palette.

A tan leather weekender bag standing upright on a bleached oak floor beside a white wall, its brass zip catching the light, twin handles folded over the top. Late afternoon light from a window on the left, a soft shadow to the right. Straight-on shot at bag height, 50mm lens, the bag centered with clear floor below. Photoreal accessories e-commerce photography, warm neutral palette.

A tan leather weekender bag standing upright on a bleached oak floor beside a white wall, its brass zip catching the light, twin handles folded over the top. Late afternoon light from a window on the left, a soft shadow to the right. Straight-on shot at bag height, 50mm lens, the bag centered with clear floor below. Photoreal accessories e-commerce photography, warm neutral palette.
Three runs of one prompt hold the brief and move the grain, the handle position and the shadow. That changes how you work: you cannot refine a render by nudging the prompt, because the nudge also rerolls everything else.
The workflow that does hold is to generate a batch, pick the frame that landed, then feed that frame back as a reference and edit it. The edit preserves what you already chose instead of gambling on it again. The editing images guide covers that step.
Tips for best results
-
Write in the five-part order. Subject, action, location, composition, style. The model places the camera after it has built the scene, and prompts written in that order need fewer retries.
-
Name the action, not just the subject. A posture or a gesture pins the figure down. Without one, the expander picks a pose and it changes every run.
-
Replace exclusions with states. There is no
negativePrompt, so "no cars" is a clause with nothing to act on. "An empty street at dawn, wet asphalt reflecting the streetlights" is a scene the model can build. -
Give light a direction and a hardness. These two words do more for consistency across a set than any style keyword.
-
Fix the subject and vary the camera when you need several shots of one product. Changing only the composition clause keeps the colorway and the finish stable across the set.
-
Generate at the tier you will deliver at. There is no upscale step in this model, and a 1K render cropped to a detail shot will not hold what a 4K render holds.
-
Batch, then edit. Generate several, choose one, and move to reference-based editing for every refinement after that.