live
MODEL IDalibaba:qwen-image@2.1-pro

Qwen-Image-2.1-Pro

Alibaba
by

Qwen-Image-2.1-Pro is the Pro tier of Alibaba's Qwen-Image-2.1 family, a unified model for text-to-image generation and prompt-guided image editing. It composes and edits from up to 10 reference images and can rewrite prompts with an LLM before generation, in a direct or agent mode with optional thinking. It outputs from 0.26 to 4.19 megapixels at aspect ratios from 1:8 to 8:1, suited to product imagery, marketing visuals, and multi-image compositions.

Qwen-Image-2.1-Pro

Composing from multiple reference images

How to combine up to ten reference images with Qwen-Image-2.1-Pro: assigning each a role, dressing a model from packshots, staging a room, and building a team photo.

Introduction

inputs.referenceImages on Qwen-Image-2.1-Pro takes up to ten images, and the slots carry no fixed meaning. The prompt assigns each image its job by position, as in "the woman from image 1" or "the trousers from image 3". One request can put people and products from separate shoots into a single frame, in a setting none of them were photographed in.

The lookbook shot below started as five studio images, one model and four products:

Built from 5 references
  • Image 1: model
  • Image 2: jacket
  • Image 3: trousers
  • Image 4: boots
  • Image 5: bag

None of the four products was ever worn in its source image, and the woman in image 1 was never photographed outdoors. For a fashion store, that is a styled outfit shot built from the packshots you already have.

Request shape

References go in inputs.referenceImages as URLs, UUIDs, or data URIs. Their order only matters because the prompt refers to it:

Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'alibaba:qwen-image@2.1-pro',
  positivePrompt: 'Dress the woman from image 1 in the outfit from the other images: the light-wash denim jacket from image 2 worn open over her black tank top, the cream wide-leg trousers from image 3, the black lug-sole ankle boots from image 4, and the burgundy shoulder bag from image 5 on her left shoulder. Keep her face, hair, and body proportions from image 1, and keep each garment\'s exact color, cut, and details from its reference. She stands in a relaxed three-quarter pose on a city sidewalk in front of a pale limestone storefront, in soft overcast daylight. Full-length fashion lookbook photography, photoreal, no text.',
  width: 1792,
  height: 2240,
  inputs: {
    referenceImages: [
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/3335fc6a-8c26-47ad-a090-fafdc06e646c.jpg',
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/5fc37fcc-ddd6-4a02-b10b-16f47afea5a8.jpg',
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/a332ee7d-dcc2-4b2d-be74-ee09b6bd02f3.jpg',
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/1d81ae8b-f659-49b1-ba2a-7fcb18e6b962.jpg',
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/1196ab82-4cdf-403a-957d-e390e04490ff.jpg'
    ]
  },
  settings: {
    promptExtend: false
  }
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "alibaba:qwen-image@2.1-pro",
            "positivePrompt": "Dress the woman from image 1 in the outfit from the other images: the light-wash denim jacket from image 2 worn open over her black tank top, the cream wide-leg trousers from image 3, the black lug-sole ankle boots from image 4, and the burgundy shoulder bag from image 5 on her left shoulder. Keep her face, hair, and body proportions from image 1, and keep each garment's exact color, cut, and details from its reference. She stands in a relaxed three-quarter pose on a city sidewalk in front of a pale limestone storefront, in soft overcast daylight. Full-length fashion lookbook photography, photoreal, no text.",
            "width": 1792,
            "height": 2240,
            "inputs": {
                "referenceImages": [
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/3335fc6a-8c26-47ad-a090-fafdc06e646c.jpg",
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/5fc37fcc-ddd6-4a02-b10b-16f47afea5a8.jpg",
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/a332ee7d-dcc2-4b2d-be74-ee09b6bd02f3.jpg",
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/1d81ae8b-f659-49b1-ba2a-7fcb18e6b962.jpg",
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/1196ab82-4cdf-403a-957d-e390e04490ff.jpg"
                ]
            },
            "settings": {
                "promptExtend": False
            }
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "imageInference",
      "taskUUID": "fbc67851-7cf2-4212-bdfb-b8b0412949d6",
      "model": "alibaba:qwen-image@2.1-pro",
      "positivePrompt": "Dress the woman from image 1 in the outfit from the other images: the light-wash denim jacket from image 2 worn open over her black tank top, the cream wide-leg trousers from image 3, the black lug-sole ankle boots from image 4, and the burgundy shoulder bag from image 5 on her left shoulder. Keep her face, hair, and body proportions from image 1, and keep each garment's exact color, cut, and details from its reference. She stands in a relaxed three-quarter pose on a city sidewalk in front of a pale limestone storefront, in soft overcast daylight. Full-length fashion lookbook photography, photoreal, no text.",
      "width": 1792,
      "height": 2240,
      "inputs": {
        "referenceImages": [
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/3335fc6a-8c26-47ad-a090-fafdc06e646c.jpg",
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/5fc37fcc-ddd6-4a02-b10b-16f47afea5a8.jpg",
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/a332ee7d-dcc2-4b2d-be74-ee09b6bd02f3.jpg",
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/1d81ae8b-f659-49b1-ba2a-7fcb18e6b962.jpg",
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/1196ab82-4cdf-403a-957d-e390e04490ff.jpg"
        ]
      },
      "settings": {
        "promptExtend": false
      }
    }
  ]'
runware run alibaba:qwen-image@2.1-pro \
  positivePrompt="Dress the woman from image 1 in the outfit from the other images: the light-wash denim jacket from image 2 worn open over her black tank top, the cream wide-leg trousers from image 3, the black lug-sole ankle boots from image 4, and the burgundy shoulder bag from image 5 on her left shoulder. Keep her face, hair, and body proportions from image 1, and keep each garment's exact color, cut, and details from its reference. She stands in a relaxed three-quarter pose on a city sidewalk in front of a pale limestone storefront, in soft overcast daylight. Full-length fashion lookbook photography, photoreal, no text." \
  width=1792 \
  height=2240 \
  inputs.referenceImages.0=https://im.runware.ai/image/os/a14d18/ws/2/ii/3335fc6a-8c26-47ad-a090-fafdc06e646c.jpg \
  inputs.referenceImages.1=https://im.runware.ai/image/os/a14d18/ws/2/ii/5fc37fcc-ddd6-4a02-b10b-16f47afea5a8.jpg \
  inputs.referenceImages.2=https://im.runware.ai/image/os/a14d18/ws/2/ii/a332ee7d-dcc2-4b2d-be74-ee09b6bd02f3.jpg \
  inputs.referenceImages.3=https://im.runware.ai/image/os/a14d18/ws/2/ii/1d81ae8b-f659-49b1-ba2a-7fcb18e6b962.jpg \
  inputs.referenceImages.4=https://im.runware.ai/image/os/a14d18/ws/2/ii/1196ab82-4cdf-403a-957d-e390e04490ff.jpg \
  settings.promptExtend=false
{
  "taskType": "imageInference",
  "taskUUID": "fbc67851-7cf2-4212-bdfb-b8b0412949d6",
  "model": "alibaba:qwen-image@2.1-pro",
  "positivePrompt": "Dress the woman from image 1 in the outfit from the other images: the light-wash denim jacket from image 2 worn open over her black tank top, the cream wide-leg trousers from image 3, the black lug-sole ankle boots from image 4, and the burgundy shoulder bag from image 5 on her left shoulder. Keep her face, hair, and body proportions from image 1, and keep each garment's exact color, cut, and details from its reference. She stands in a relaxed three-quarter pose on a city sidewalk in front of a pale limestone storefront, in soft overcast daylight. Full-length fashion lookbook photography, photoreal, no text.",
  "width": 1792,
  "height": 2240,
  "inputs": {
    "referenceImages": [
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/3335fc6a-8c26-47ad-a090-fafdc06e646c.jpg",
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/5fc37fcc-ddd6-4a02-b10b-16f47afea5a8.jpg",
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/a332ee7d-dcc2-4b2d-be74-ee09b6bd02f3.jpg",
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/1d81ae8b-f659-49b1-ba2a-7fcb18e6b962.jpg",
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/1196ab82-4cdf-403a-957d-e390e04490ff.jpg"
    ]
  },
  "settings": {
    "promptExtend": false
  }
}
Response
[
  {
    "taskType": "imageInference",
    "taskUUID": "fbc67851-7cf2-4212-bdfb-b8b0412949d6",
    "imageUUID": "1f39b83a-6c67-411c-a2b7-974970cdd5d9",
    "imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/1f39b83a-6c67-411c-a2b7-974970cdd5d9.jpg"
  }
]

The output size follows the same pixel budget as generation, up to 4,194,304 pixels, whatever the sizes of the references. Leave width and height out and the first reference sets the shape, so set both whenever the first image is not the frame you want.

Each reference can be up to 10 MB. With references attached, promptExtendMode accepts only direct, so a request that also sends agent is rejected. Every example here sets promptExtend to false, which keeps the roles you assign in the prompt from being rewritten.

Assigning roles by position

The model has no idea which image is the person and which is the jacket until the prompt says so. The hero prompt reads as a roll call:

Dress the woman from image 1 in the outfit from the other images:
the light-wash denim jacket from image 2 worn open over her black tank top,
the cream wide-leg trousers from image 3,
the black lug-sole ankle boots from image 4,
and the burgundy shoulder bag from image 5 on her left shoulder.
Keep her face, hair, and body proportions from image 1,
and keep each garment's exact color, cut, and details from its reference.

Four habits make the roll call work:

  • Number every image once, in the order it sits in the array.
  • Pair each number with a short description, as in "the burgundy shoulder bag from image 5". The description tells the model what to look for in that image and still points at the right item if you reorder the array and forget a number.
  • Say where each item goes, as in "worn open" or "on her left shoulder".
  • Close with a keep clause that names what must survive from the references, such as the face and the colors.

Staging a room from a furniture set

One reference can be the setting and the rest the things that go in it. Here the first image is an empty apartment and the other six are catalog shots:

Built from 7 references
  • Image 1: room
  • Image 2: sofa
  • Image 3: coffee table
  • Image 4: rug
  • Image 5: lamp
  • Image 6: print
  • Image 7: olive tree

Every product has a position relative to the room or to another product, such as "against the left wall" or "in front of the sofa". A list of products with no placement leaves the floor plan to the model, and it will pick a different one on every call.

The room itself is a reference too, so its window and its light carry into the result. That is what makes the output read as the listing's own apartment furnished, which is the virtual staging a property platform sells, rather than a generic interior with the same furniture.

A team photo from individual headshots

References work for people as well as products. Four headshots taken on different days become one group photo for a company's About page:

Built from 4 references
  • Image 1
  • Image 2
  • Image 3
  • Image 4

"From left to right" followed by the image numbers fixes the lineup, so the order on the page matches the order you chose. The headshots only show head and shoulders, which means everything below the collar in the group shot was invented to match the clothing each person wore.

Check every face against its source before a result like this goes live. A team page is a photo of real colleagues, and a face that drifted toward a stranger is worse than no group shot at all.

Filling all ten slots

Ten references is the ceiling. This flat-lay uses all of them, one per product in an outdoor retailer's camping bundle:

Built from 10 references
  • 1: tent
  • 2: lantern
  • 3: stove
  • 4: mug
  • 5: sleeping bag
  • 6: headlamp
  • 7: bottle
  • 8: knife
  • 9: cook pot
  • 10: boots

With ten items the roll call gets long, so each product gets a short name ("the lantern from image 2"), with a color only where two items could be mistaken for each other. The references already hold the materials, and the prompt spends its words on the layout instead.

Give the layout exactly as many slots as you have products. "Two rows of five" leaves no empty place for the model to fill, and a loose grid does: asked for one, it returned eleven objects from ten references, with one product doubled.

Count the output against the inputs at this size. Ten products are easy to check by eye, and a bundle shot with a missing or doubled item is a listing that misrepresents what ships.

Tips

  1. Refer to every image by number and description. The number says which slot, and the description says what to take from it.

  2. Place each item relative to something already in the frame. "In front of the sofa" is a position, and "in the room" leaves the layout to the model.

  3. Close with a keep clause. Faces and colors survive when the prompt says they must.

  4. Check the output against the references. Count the products and compare every face before a composite goes live.