live
MODEL IDalibaba:qwen-image@2.1-pro

Qwen-Image-2.1-Pro

Alibaba
by

Qwen-Image-2.1-Pro is the Pro tier of Alibaba's Qwen-Image-2.1 family, a unified model for text-to-image generation and prompt-guided image editing. It composes and edits from up to 10 reference images and can rewrite prompts with an LLM before generation, in a direct or agent mode with optional thinking. It outputs from 0.26 to 4.19 megapixels at aspect ratios from 1:8 to 8:1, suited to product imagery, marketing visuals, and multi-image compositions.

Qwen-Image-2.1-Pro

Editing images

How to edit a single image with Qwen-Image-2.1-Pro: naming what changes and what holds, swapping label copy, circling the region to edit, and chaining edits.

Introduction

Pass one image in inputs.referenceImages with an instruction in positivePrompt, and Qwen-Image-2.1-Pro returns the same image with that change applied. There is no mask field and no edit mode to pick. The instruction is the whole interface, so its wording is the thing to get right.

A white leather low-top sneaker in side profile with a perforated toe and a white cupsole on a light gray backdrop

A product photograph of a low-top white leather sneaker in side profile facing right, with white laces, a perforated toe box, a tonal white heel tab, and a flat white rubber cupsole. Standing on a plain light gray studio backdrop with a soft contact shadow. Even soft studio lighting, the shoe centered with clear space on both sides. Photoreal footwear e-commerce photography, no logos, no text.

SourceEdited

Drag the slider across the toe box. The perforations and the laces hold their place, because the instruction, shown in the request below, listed them as fixed. One packshot becomes a second colorway without a second shoot.

This guide covers scoping an instruction, editing the copy on packaging, pointing at a region by drawing on the source, and chaining several edits on one image.

Request shape

An edit is a generation with one reference attached:

Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'alibaba:qwen-image@2.1-pro',
  positivePrompt: 'Change the sneaker\'s upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged.',
  width: 2496,
  height: 1664,
  inputs: {
    referenceImages: [
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg'
    ]
  },
  settings: {
    promptExtend: false
  }
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "alibaba:qwen-image@2.1-pro",
            "positivePrompt": "Change the sneaker's upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged.",
            "width": 2496,
            "height": 1664,
            "inputs": {
                "referenceImages": [
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg"
                ]
            },
            "settings": {
                "promptExtend": False
            }
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "imageInference",
      "taskUUID": "f3c7f7af-6c5b-41a1-a7d8-05c9f6e7b702",
      "model": "alibaba:qwen-image@2.1-pro",
      "positivePrompt": "Change the sneaker's upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged.",
      "width": 2496,
      "height": 1664,
      "inputs": {
        "referenceImages": [
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg"
        ]
      },
      "settings": {
        "promptExtend": false
      }
    }
  ]'
runware run alibaba:qwen-image@2.1-pro \
  positivePrompt="Change the sneaker's upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged." \
  width=2496 \
  height=1664 \
  inputs.referenceImages.0=https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg \
  settings.promptExtend=false
{
  "taskType": "imageInference",
  "taskUUID": "f3c7f7af-6c5b-41a1-a7d8-05c9f6e7b702",
  "model": "alibaba:qwen-image@2.1-pro",
  "positivePrompt": "Change the sneaker's upper from white leather to forest green suede and the sole from white rubber to a honey gum rubber. Keep the white laces, the perforations on the toe box, the exact shape and side profile of the shoe, the light gray backdrop, the contact shadow, and the lighting unchanged.",
  "width": 2496,
  "height": 1664,
  "inputs": {
    "referenceImages": [
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/36b493c2-14d9-428e-be60-2fedc9406e54.jpg"
    ]
  },
  "settings": {
    "promptExtend": false
  }
}
Response
[
  {
    "taskType": "imageInference",
    "taskUUID": "f3c7f7af-6c5b-41a1-a7d8-05c9f6e7b702",
    "imageUUID": "8d5c357a-3c32-4401-be0f-42f1c7f24e59",
    "imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/8d5c357a-3c32-4401-be0f-42f1c7f24e59.jpg"
  }
]

Two fields need attention. Leave width and height out and the output takes the shape of the reference, scaled down when its longer side is above 2048 pixels. This source is 2496 × 1664, so the request sets both dimensions and the edit drops in where the original was. And promptExtend is on by default, which hands your instruction to an LLM to rewrite before the model reads it. Every example here turns it off, so the instruction you read is the one that ran. Prompt extension shows the same kind of edit with it on.

Naming what changes and what holds

An edit instruction has two halves: the change and the list of what stays. The first half is what you want. The second half is what stops the model from reinterpreting everything else in the frame:

Paint the back wall a deep sage green in a matte finish. Keep the sofa, the cushions, the coffee table, the rug, the plant, the window, the curtains, the floor, the lighting, and the camera position exactly as they are.
A bright living room with a light gray sofa and ochre cushions against a white wall, an oak coffee table on a cream rug, a fiddle-leaf fig, and a window with sheer curtains

A real estate listing photograph of a bright living room. A low light gray fabric sofa with two ochre cushions sits against a plain white back wall, a round oak coffee table on a cream wool rug in front of it, a tall fiddle-leaf fig in a white pot to the left, and a large window on the right with sheer white curtains. Pale oak floorboards. Soft natural daylight from the window, wide shot from the doorway at chest height. Photoreal interiors photography, no people, no text.

SourceRepainted

The second sentence is a roll call of the room, and everything on it came through in place. For a listing or a paint retailer, that is the difference between a color preview and a different room: the buyer has to recognize the space they toured.

Write the keep list from what is in the frame, not from a generic template. "Keep everything else the same" names nothing, while "the sofa, the cushions, the coffee table" gives the model a checklist to hold against.

Changing copy on packaging

Text on a product is edited the same way as color or material. Quote the old string and the new one, and name the copy that must not move:

Change the flavor name on the can from "LEMON LIME" to "BLOOD ORANGE", and change the can's color from pale lime green to a deep coral with the flavor name in dark red. Keep the brand name "SOLÉA" in the same white typeface, size, and position, and keep the white wave pattern, the can's shape, the condensation, the lighting, and the white backdrop unchanged.
A slim pale lime green can printed with the brand name SOLÉA in white and the flavor LEMON LIME in dark green, with a white wave pattern at the base

A product photograph of a slim 330 ml aluminum can of sparkling water standing upright on a plain white studio backdrop. The can is printed in pale lime green with a white wave pattern around its base. The brand name "SOLÉA" is printed across the upper half in a bold white sans serif, and below it the flavor name "LEMON LIME" in smaller dark green capitals. Condensation droplets on the metal, soft studio lighting with a bright vertical highlight down the left side. Photoreal beverage packshot, centered.

LEMON LIMEBLOOD ORANGE

The brand name kept its typeface and its place on the can, accent included. That is one flavor's packshot becoming the whole range: render the first can properly, then edit the name and the color for each variant.

Quoting the old string matters as much as quoting the new one. "Change the flavor name" alone leaves the model to decide which of the two lines on the can is the flavor.

Pointing at a region

When the target is one object among several, draw a circle around it on the source and refer to the circle in the instruction. The mark does the targeting, so the wording only has to describe the replacement:

Replace the object inside the red circle with a ribbed amber glass reed diffuser holding a bundle of thin black reeds, standing on the sideboard with a soft contact shadow. Remove the red circle. Keep everything outside the circle exactly as it is.
A walnut sideboard with a brass lamp on the left, a stack of art books in the middle, and a white vase with pampas stems on the right circled in red, under a framed abstract print

A lifestyle photograph for a home decor store: a low walnut sideboard against a plain warm white wall. On the left end of the sideboard, a brass table lamp with a white linen shade. In the middle, a short stack of three hardcover art books. On the right end, a tall white ceramic vase holding a few dried pampas stems. A framed abstract print in muted earth tones hangs above the center of the sideboard. Soft afternoon daylight from the left, straight-on view at sideboard height. Photoreal interiors photography, no people, no text.

Marked sourceEdited

The instruction never says "the vase". It says "the object inside the red circle", so the same sentence works for any object in the frame. That suits a pipeline where a user taps the item to swap and your code draws the circle.

Draw the mark into the image you send, in a color that appears nowhere else in the frame, and ask for the mark to be removed as part of the same instruction.

Chaining edits

Each edit returns an ordinary image, so the output of one round can be the reference for the next. The chain below restyles one catalog shot in six rounds, one change per round:

Step through the rounds, and go back to Original to compare. His face and his pose hold through all six, because every instruction restated them. The keep list does not carry over from one round to the next: each edit only sees its reference image and its own prompt, so anything a round fails to name is open to change.

What a round costs

An edit redraws the whole frame, including everything the instruction told it to keep. Each pass leaves a little more grain behind, and the inset on each image follows it across his face. Measured on the plain backdrop, as the standard deviation of three flat patches on a 0 to 255 scale:

ImageBackdrop grain
Original1.7
Round 12.5
Round 23.6
Round 34.0
Round 44.0
Round 55.4
Round 64.3

Six rounds leave about two and a half times the grain of the original. The same restyle also fits in one instruction applied to the original, which pays for a single pass:

Fold the changes you already know into one instruction, and chain only the ones that depend on seeing a result first. When a chain has grown long, go back to the original and replay the approved changes in a single request.

Tips

  1. Set the output size when the source is large. Without width and height the result follows the source's shape, capped at 2048 pixels on the longer side.

  2. Name what holds, not only what changes. The keep list is what stops the model from rebuilding the rest of the frame.

  3. Quote the old string and the new one in copy edits. Both quotes tell the model which line of text to replace and exactly what to write.

  4. Circle the target when the scene has look-alikes. A red mark on the source is more precise than a description of where the object sits.

  5. Keep extension off. With promptExtend set to false, the model reads the instruction you wrote.

  6. Fold known changes into one instruction. Every round redraws the frame and adds grain, so chain only the edits that depend on the previous result.