live
MODEL IDalibaba:qwen-image@2.1-pro

Qwen-Image-2.1-Pro

Alibaba
by

Qwen-Image-2.1-Pro is the Pro tier of Alibaba's Qwen-Image-2.1 family, a unified model for text-to-image generation and prompt-guided image editing. It composes and edits from up to 10 reference images and can rewrite prompts with an LLM before generation, in a direct or agent mode with optional thinking. It outputs from 0.26 to 4.19 megapixels at aspect ratios from 1:8 to 8:1, suited to product imagery, marketing visuals, and multi-image compositions.

Qwen-Image-2.1-Pro

Prompting

How to prompt Qwen-Image-2.1-Pro, from sizing inside its pixel budget and banner ratios to layering a scene, quoting on-image copy, and locking a result with a seed.

Introduction

Qwen-Image-2.1-Pro is the Pro tier of Alibaba's Qwen-Image-2.1 family, and it is one model for generating and editing. Send a prompt on its own and it generates. Attach reference images and it edits or composes from them.

Two things shape how you write for it. Output size is a pixel budget with free aspect ratios, so you set the exact shape a layout needs. And an LLM rewrites your prompt before generation unless you turn that off.

Every example in this guide sets promptExtend to false, so the prompt shown with each image is the prompt the model received. The rewrite has its own guide. Working from reference images is covered in Editing images and Composing from multiple references.

Request shape

A generation needs a positivePrompt of 2 to 3,000 characters. width and height are optional, but they default to 512 × 512, the smallest square the model accepts, so set them on every request you intend to keep.

Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'alibaba:qwen-image@2.1-pro',
  positivePrompt: 'A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.',
  width: 2496,
  height: 1664,
  settings: {
    promptExtend: false
  }
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "alibaba:qwen-image@2.1-pro",
            "positivePrompt": "A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.",
            "width": 2496,
            "height": 1664,
            "settings": {
                "promptExtend": False
            }
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "imageInference",
      "taskUUID": "4a84bed8-9c57-4af8-943c-601cbb223579",
      "model": "alibaba:qwen-image@2.1-pro",
      "positivePrompt": "A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.",
      "width": 2496,
      "height": 1664,
      "settings": {
        "promptExtend": false
      }
    }
  ]'
runware run alibaba:qwen-image@2.1-pro \
  positivePrompt="A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text." \
  width=2496 \
  height=1664 \
  settings.promptExtend=false
{
  "taskType": "imageInference",
  "taskUUID": "4a84bed8-9c57-4af8-943c-601cbb223579",
  "model": "alibaba:qwen-image@2.1-pro",
  "positivePrompt": "A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.",
  "width": 2496,
  "height": 1664,
  "settings": {
    "promptExtend": false
  }
}
Response
[
  {
    "taskType": "imageInference",
    "taskUUID": "4a84bed8-9c57-4af8-943c-601cbb223579",
    "imageUUID": "b62d8137-9ae4-4537-8bdb-c5d7dc4c17aa",
    "imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/b62d8137-9ae4-4537-8bdb-c5d7dc4c17aa.jpg"
  }
]

Sizing inside the pixel budget

Any width and height are valid as long as their product lands between 262,144 and 4,194,304 pixels and the ratio stays between 1:8 and 8:1. Neither side can be longer than 4096 pixels, which is what caps the widest shapes. The ceiling is 2048 × 2048 as a square, and the same budget stretches into other shapes:

ShapeSizePixels
1:12048 × 20484,194,304
3:22496 × 16644,153,344
4:51792 × 22404,014,080
16:92560 × 14403,686,400
4:14096 × 10244,194,304
8:14096 × 5122,097,152

Reference images do not shrink the budget, so a size that works for generation also works for an edit.

The wide end of the range reaches web banner formats in a single pass, with no larger frame to generate and crop. An 8:1 strip is the shape of a site header or a leaderboard ad slot:

A 4:1 frame fits a homepage hero with room for a headline:

Both prompts describe the scene left to right, naming what sits at each end and what sits in the center. A frame that wide has room for several subjects side by side, and without an order in the prompt the model picks one for you. Copy space is part of that order: the furniture hero assigns its left third to empty wall, which is where the page's headline goes.

Layering a prompt

A short prompt is a valid prompt. It just hands every decision to the model:

Three words returned a bowl, with the ingredients and the camera angle chosen for you. A delivery app needs every dish on the menu shot the same way, and that consistency only comes from writing the choices down.

The longer prompt is built in layers:

A salmon poke bowl for a food delivery app menu, shot from directly overhead, A matte black ceramic bowl sits centered on a pale ash wood table, Inside, sushi rice fills the base, topped in neat wedges with cubes of raw salmon, sliced avocado fanned in a crescent, shelled edamame, shredded purple cabbage, and pickled ginger, finished with black sesame seeds and a drizzle of spicy mayo, A pair of light wooden chopsticks rests across the right edge of the bowl, Soft diffused daylight from the upper left, gentle shadow to the lower right, Photoreal food photography, vivid natural color, no text, no hands
SubjectCameraCompositionDetailPropsLightingStyle

Lead with the subject and its purpose. "For a food delivery app menu" tells the model what kind of photograph this is before any styling arrives. Name positions and counts ("centered", "across the right edge", "a pair"), since anything left vague is something the model will decide on its own.

Quoting on-image copy

Wrap every string you want rendered in double quotes, then say where it sits and how it is set. The quotes separate the words to print from the words that describe the image:

Every quoted string came through character for character, including the en dash in the poster's dates and all four characters of the Chinese headline. Place each line next to its quote: the poster's three levels of copy are three separate instructions, each with a size, a weight, and a position, and a line you quote without placing lands wherever the model finds room.

Write the words in the case you want printed. "RADIANCE SERUM" was typed in capitals, so the capitals are not left to the typeface.

Seeds and repeatability

A fixed value pins the noise the render starts from. The same prompt with the same seed returns the same image, and a different seed returns another take on the same description:

The first two calls returned the same file, byte for byte. The third keeps every item the prompt named and changes what it left open, such as the cut of the coat and where she looks. A seed is how you keep a render you liked while you adjust the wording around it, and how you hand a colleague the exact image instead of a description of it.

Reproducibility needs promptExtend set to false. With extension on, the same seed and prompt return a different image on every call, because the rewrite that runs first is not seeded. Prompt extension shows the difference.

Tips

  1. Set both dimensions on every request. Leaving width and height out returns 512 × 512, a fraction of what the budget allows.

  2. Turn extension off while you iterate. With promptExtend on you are tuning a prompt the model never sees, and two runs of the same wording will not match.

  3. Pick the shape before the wording. A strip and a portrait frame need different composition language, and the words only pay off if the canvas has room for them.

  4. Quote every string that has to be spelled exactly. Write it in the case you want printed, and describe its size and position alongside it.

  5. Keep a seed with every render you might need again. It costs nothing to send and it is the only way back to a specific image.