Qwen-Image-2.1-Pro

Qwen-Image-2.1-Pro is the Pro tier of Alibaba's Qwen-Image-2.1 family, a unified model for text-to-image generation and prompt-guided image editing. It composes and edits from up to 10 reference images and can rewrite prompts with an LLM before generation, in a direct or agent mode with optional thinking. It outputs from 0.26 to 4.19 megapixels at aspect ratios from 1:8 to 8:1, suited to product imagery, marketing visuals, and multi-image compositions.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
← All GuidesEditing images
How to edit a single image with Qwen-Image-2.1-Pro: naming what changes and what holds, swapping label copy, circling the region to edit, and chaining edits.
Introduction
Pass one image in inputs.referenceImages with an instruction in positivePrompt, and Qwen-Image-2.1-Pro returns the same image with that change applied. There is no mask field and no edit mode to pick. The instruction is the whole interface, so its wording is the thing to get right.
Drag the slider across the toe box. The perforations and the laces hold their place, because the instruction, shown in the request below, listed them as fixed. One packshot becomes a second colorway without a second shoot.
This guide covers scoping an instruction, editing the copy on packaging, pointing at a region by drawing on the source, and chaining several edits on one image.
Request shape
An edit is a generation with one reference attached:
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'alibaba:qwen-image@2.1-pro',
positivePrompt: 'Change the sneaker\'s upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged.',
width: 2496,
height: 1664,
inputs: {
referenceImages: [
'https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg'
]
},
settings: {
promptExtend: false
}
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "alibaba:qwen-image@2.1-pro",
"positivePrompt": "Change the sneaker's upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged.",
"width": 2496,
"height": 1664,
"inputs": {
"referenceImages": [
"https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg"
]
},
"settings": {
"promptExtend": False
}
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "imageInference",
"taskUUID": "f3c7f7af-6c5b-41a1-a7d8-05c9f6e7b702",
"model": "alibaba:qwen-image@2.1-pro",
"positivePrompt": "Change the sneaker's upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged.",
"width": 2496,
"height": 1664,
"inputs": {
"referenceImages": [
"https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg"
]
},
"settings": {
"promptExtend": false
}
}
]'runware run alibaba:qwen-image@2.1-pro \
positivePrompt="Change the sneaker's upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged." \
width=2496 \
height=1664 \
inputs.referenceImages.0=https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg \
settings.promptExtend=false{
"taskType": "imageInference",
"taskUUID": "f3c7f7af-6c5b-41a1-a7d8-05c9f6e7b702",
"model": "alibaba:qwen-image@2.1-pro",
"positivePrompt": "Change the sneaker's upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged.",
"width": 2496,
"height": 1664,
"inputs": {
"referenceImages": [
"https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg"
]
},
"settings": {
"promptExtend": false
}
}Response
[
{
"taskType": "imageInference",
"taskUUID": "f3c7f7af-6c5b-41a1-a7d8-05c9f6e7b702",
"imageUUID": "8d5c357a-3c32-4401-be0f-42f1c7f24e59",
"imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/8d5c357a-3c32-4401-be0f-42f1c7f24e59.jpg"
}
]Two fields need attention. Leave width and height out and the output takes the shape of the reference, scaled down when its longer side is above 2048 pixels. This source is 2496 × 1664, so the request sets both dimensions and the edit drops in where the original was. And promptExtend is on by default, which hands your instruction to an LLM to rewrite before the model reads it. Every example here turns it off, so the instruction you read is the one that ran. Prompt extension shows the same kind of edit with it on.
Naming what changes and what holds
An edit instruction has two halves: the change and the list of what stays. The first half is what you want. The second half is what stops the model from reinterpreting everything else in the frame:
Paint the back wall a deep sage green in a matte finish. Keep the sofa, the cushions, the coffee table, the rug, the plant, the window, the curtains, the floor, the lighting, and the camera position exactly as they are.The second sentence is a roll call of the room, and everything on it came through in place. For a listing or a paint retailer, that is the difference between a color preview and a different room: the buyer has to recognize the space they toured.
Write the keep list from what is in the frame, not from a generic template. "Keep everything else the same" names nothing, while "the sofa, the cushions, the coffee table" gives the model a checklist to hold against.
Changing copy on packaging
Text on a product is edited the same way as color or material. Quote the old string and the new one, and name the copy that must not move:
Change the flavor name on the can from "LEMON LIME" to "BLOOD ORANGE", and change the can's color from pale lime green to a deep coral with the flavor name in dark red. Keep the brand name "SOLÉA" in the same white typeface, size, and position, and keep the white wave pattern, the can's shape, the condensation, the lighting, and the white backdrop unchanged.The brand name kept its typeface and its place on the can, accent included. That is one flavor's packshot becoming the whole range: render the first can properly, then edit the name and the color for each variant.
Quoting the old string matters as much as quoting the new one. "Change the flavor name" alone leaves the model to decide which of the two lines on the can is the flavor.
Pointing at a region
When the target is one object among several, draw a circle around it on the source and refer to the circle in the instruction. The mark does the targeting, so the wording only has to describe the replacement:
Replace the object inside the red circle with a ribbed amber glass reed diffuser holding a bundle of thin black reeds, standing on the sideboard with a soft contact shadow. Remove the red circle. Keep everything outside the circle exactly as it is.The instruction never says "the vase". It says "the object inside the red circle", so the same sentence works for any object in the frame. That suits a pipeline where a user taps the item to swap and your code draws the circle.
Draw the mark into the image you send, in a color that appears nowhere else in the frame, and ask for the mark to be removed as part of the same instruction.
Chaining edits
Each edit returns an ordinary image, so the output of one round can be the reference for the next. The chain below restyles one catalog shot in six rounds, one change per round:

An e-commerce catalog photo of a man in his early thirties against a plain light gray studio backdrop. Short dark hair, trimmed beard, relaxed expression. He stands in an easy three-quarter stance with his weight on one leg and one hand in his jeans pocket. He wears a plain white crew-neck T-shirt, straight-leg mid-blue jeans, and white canvas sneakers. Soft even studio lighting from the front, subtle floor shadow. Photoreal fashion catalog photography, full length, centered, no text.

Replace his white T-shirt with a navy and white horizontal striped long-sleeve Breton top. Keep his face, hair, pose, jeans, sneakers, the backdrop, and the lighting unchanged.

Add an unbuttoned tan cotton trench coat over his striped top, falling to just above the knee. Keep his face, hair, pose, striped top, jeans, sneakers, the backdrop, and the lighting unchanged.

Change the trench coat from tan to olive green. Keep its cut, its buttons, and its length, and keep his face, hair, pose, striped top, jeans, sneakers, the backdrop, and the lighting unchanged.

Replace his blue jeans with charcoal gray wool trousers with a straight leg. Keep his face, hair, pose, olive trench coat, striped top, sneakers, the backdrop, and the lighting unchanged.

Replace his white sneakers with brown leather ankle boots. Keep his face, hair, pose, olive trench coat, striped top, charcoal trousers, the backdrop, and the lighting unchanged.

Add a light gray wool scarf looped once around his neck. Keep his face, hair, pose, olive trench coat, striped top, charcoal trousers, boots, the backdrop, and the lighting unchanged.
Step through the rounds, and go back to Original to compare. His face and his pose hold through all six, because every instruction restated them. The keep list does not carry over from one round to the next: each edit only sees its reference image and its own prompt, so anything a round fails to name is open to change.
What a round costs
An edit redraws the whole frame, including everything the instruction told it to keep. Each pass leaves a little more grain behind, and the inset on each image follows it across his face. Measured on the plain backdrop, as the standard deviation of three flat patches on a 0 to 255 scale:
| Image | Backdrop grain |
|---|---|
| Original | 1.7 |
| Round 1 | 2.5 |
| Round 2 | 3.6 |
| Round 3 | 4.0 |
| Round 4 | 4.0 |
| Round 5 | 5.4 |
| Round 6 | 4.3 |
Six rounds leave about two and a half times the grain of the original. The same restyle also fits in one instruction applied to the original, which pays for a single pass:

Add a light gray wool scarf looped once around his neck. Keep his face, hair, pose, olive trench coat, striped top, charcoal trousers, boots, the backdrop, and the lighting unchanged.

Replace his white T-shirt with a navy and white horizontal striped long-sleeve Breton top, and add an unbuttoned olive green cotton trench coat over it, falling to just above the knee. Replace his blue jeans with charcoal gray wool trousers with a straight leg and his white sneakers with brown leather ankle boots, and add a light gray wool scarf looped once around his neck. Keep his face, hair, pose, the backdrop, and the lighting unchanged.
Fold the changes you already know into one instruction, and chain only the ones that depend on seeing a result first. When a chain has grown long, go back to the original and replay the approved changes in a single request.
Tips
-
Set the output size when the source is large. Without
widthandheightthe result follows the source's shape, capped at 2048 pixels on the longer side. -
Name what holds, not only what changes. The keep list is what stops the model from rebuilding the rest of the frame.
-
Quote the old string and the new one in copy edits. Both quotes tell the model which line of text to replace and exactly what to write.
-
Circle the target when the scene has look-alikes. A red mark on the source is more precise than a description of where the object sits.
-
Keep extension off. With
promptExtendset tofalse, the model reads the instruction you wrote. -
Fold known changes into one instruction. Every round redraws the frame and adds grain, so chain only the edits that depend on the previous result.







