Qwen-Image-2.1-Pro

Qwen-Image-2.1-Pro is the Pro tier of Alibaba's Qwen-Image-2.1 family, a unified model for text-to-image generation and prompt-guided image editing. It composes and edits from up to 10 reference images and can rewrite prompts with an LLM before generation, in a direct or agent mode with optional thinking. It outputs from 0.26 to 4.19 megapixels at aspect ratios from 1:8 to 8:1, suited to product imagery, marketing visuals, and multi-image compositions.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
← All GuidesPrompting
How to prompt Qwen-Image-2.1-Pro, from sizing inside its pixel budget and banner ratios to layering a scene, quoting on-image copy, and locking a result with a seed.
Introduction
Qwen-Image-2.1-Pro is the Pro tier of Alibaba's Qwen-Image-2.1 family, and it is one model for generating and editing. Send a prompt on its own and it generates. Attach reference images and it edits or composes from them.
Two things shape how you write for it. Output size is a pixel budget with free aspect ratios, so you set the exact shape a layout needs. And an LLM rewrites your prompt before generation unless you turn that off.

A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.
Every example in this guide sets promptExtend to false, so the prompt shown with each image is the prompt the model received. The rewrite has its own guide. Working from reference images is covered in Editing images and Composing from multiple references.
Request shape
A generation needs a positivePrompt of 2 to 3,000 characters. width and height are optional, but they default to 512 × 512, the smallest square the model accepts, so set them on every request you intend to keep.
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'alibaba:qwen-image@2.1-pro',
positivePrompt: 'A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.',
width: 2496,
height: 1664,
settings: {
promptExtend: false
}
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "alibaba:qwen-image@2.1-pro",
"positivePrompt": "A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.",
"width": 2496,
"height": 1664,
"settings": {
"promptExtend": False
}
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "imageInference",
"taskUUID": "4a84bed8-9c57-4af8-943c-601cbb223579",
"model": "alibaba:qwen-image@2.1-pro",
"positivePrompt": "A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.",
"width": 2496,
"height": 1664,
"settings": {
"promptExtend": false
}
}
]'runware run alibaba:qwen-image@2.1-pro \
positivePrompt="A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text." \
width=2496 \
height=1664 \
settings.promptExtend=false{
"taskType": "imageInference",
"taskUUID": "4a84bed8-9c57-4af8-943c-601cbb223579",
"model": "alibaba:qwen-image@2.1-pro",
"positivePrompt": "A sportswear campaign photograph of a woman in her early thirties running along a seaside promenade at sunrise, caught mid-stride from a low side angle. She wears a fitted charcoal long-sleeve running top, black running tights, and white running shoes with a coral sole. The pale stone paving and a low white railing lead toward a distant pier, with a calm sea and a pale gold sky behind her. Warm low sun from the right rims her hair and shoulders. Photoreal athletic advertising photography, fast shutter, crisp motion, clean negative space on the left, no logos, no text.",
"width": 2496,
"height": 1664,
"settings": {
"promptExtend": false
}
}Response
[
{
"taskType": "imageInference",
"taskUUID": "4a84bed8-9c57-4af8-943c-601cbb223579",
"imageUUID": "b62d8137-9ae4-4537-8bdb-c5d7dc4c17aa",
"imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/b62d8137-9ae4-4537-8bdb-c5d7dc4c17aa.jpg"
}
]Sizing inside the pixel budget
Any width and height are valid as long as their product lands between 262,144 and 4,194,304 pixels and the ratio stays between 1:8 and 8:1. Neither side can be longer than 4096 pixels, which is what caps the widest shapes. The ceiling is 2048 × 2048 as a square, and the same budget stretches into other shapes:
| Shape | Size | Pixels |
|---|---|---|
| 1:1 | 2048 × 2048 | 4,194,304 |
| 3:2 | 2496 × 1664 | 4,153,344 |
| 4:5 | 1792 × 2240 | 4,014,080 |
| 16:9 | 2560 × 1440 | 3,686,400 |
| 4:1 | 4096 × 1024 | 4,194,304 |
| 8:1 | 4096 × 512 | 2,097,152 |
Reference images do not shrink the budget, so a size that works for generation also works for an edit.
The wide end of the range reaches web banner formats in a single pass, with no larger frame to generate and crop. An 8:1 strip is the shape of a site header or a leaderboard ad slot:

A panoramic website header strip for an alpine ski resort, laid out left to right. On the far left, a timber mountain lodge with glowing windows and snow on its roof. In the center, a wide groomed piste with a few distant skiers carving turns. On the right, a chairlift climbing toward a sunlit ridge, and on the far right, jagged snowy peaks against a clear deep blue sky. Late afternoon sun from the right, clean snow in the foreground with no people and no shadows of people. Photoreal travel photography, crisp and bright, no text, no logos.
A 4:1 frame fits a homepage hero with room for a headline:

A wide homepage hero for an online furniture store: a long, bright living room laid out left to right. The left third of the frame is empty warm white wall, kept clear for a headline. In the center, a low three-seat sofa in terracotta linen with two cream cushions, behind a round travertine coffee table on a large jute rug. On the right, a pale oak sideboard with a trailing plant, and at the far right edge a tall window. Pale oak floorboards, soft morning daylight from the window. Photoreal interiors photography, eye-level straight-on view, calm warm palette, no people, no text.
Both prompts describe the scene left to right, naming what sits at each end and what sits in the center. A frame that wide has room for several subjects side by side, and without an order in the prompt the model picks one for you. Copy space is part of that order: the furniture hero assigns its left third to empty wall, which is where the page's headline goes.
Layering a prompt
A short prompt is a valid prompt. It just hands every decision to the model:

a poke bowl

A salmon poke bowl for a food delivery app menu, shot from directly overhead. A matte black ceramic bowl sits centered on a pale ash wood table. Inside, sushi rice fills the base, topped in neat wedges with cubes of raw salmon, sliced avocado fanned in a crescent, shelled edamame, shredded purple cabbage, and pickled ginger, finished with black sesame seeds and a drizzle of spicy mayo. A pair of light wooden chopsticks rests across the right edge of the bowl. Soft diffused daylight from the upper left, gentle shadow to the lower right. Photoreal food photography, vivid natural color, no text, no hands.
Three words returned a bowl, with the ingredients and the camera angle chosen for you. A delivery app needs every dish on the menu shot the same way, and that consistency only comes from writing the choices down.
The longer prompt is built in layers:
Lead with the subject and its purpose. "For a food delivery app menu" tells the model what kind of photograph this is before any styling arrives. Name positions and counts ("centered", "across the right edge", "a pair"), since anything left vague is something the model will decide on its own.
Quoting on-image copy
Wrap every string you want rendered in double quotes, then say where it sits and how it is set. The quotes separate the words to print from the words that describe the image:

A printed poster for a design conference. Deep cobalt blue background with a large abstract composition of overlapping cream and orange geometric shapes filling the upper two thirds. Across the lower third, left-aligned, the headline "FORM & FUNCTION" in a heavy condensed sans serif in cream, with "Design Conference 2026" below it in a lighter weight. At the bottom, a single line in small caps reads "LISBON · 14–16 OCTOBER · TICKETS ON SALE NOW". Flat graphic design, crisp print-ready typography, generous margins.

A social media launch post for a skincare brand. A plain frosted glass serum bottle with a gold dropper and no label stands on a pale pink stone plinth against a warm blush background, with a soft shadow behind it. Above the bottle, centered, the Chinese headline "焕亮精华" in an elegant thin serif, with the English line "RADIANCE SERUM" directly beneath it in small, widely spaced capitals. In the lower left corner, a rounded white tag reads "新品上市" with "NEW ARRIVAL" beneath it in small capitals. No other text anywhere in the image. Clean beauty advertising layout, soft studio light, photoreal product, crisp typography.
Every quoted string came through character for character, including the en dash in the poster's dates and all four characters of the Chinese headline. Place each line next to its quote: the poster's three levels of copy are three separate instructions, each with a size, a weight, and a position, and a line you quote without placing lands wherever the model finds room.
Write the words in the case you want printed. "RADIANCE SERUM" was typed in capitals, so the capitals are not left to the typeface.
Seeds and repeatability
A fixed value pins the noise the render starts from. The same prompt with the same seed returns the same image, and a different seed returns another take on the same description:

An e-commerce catalog photo of a woman in her late twenties standing in a relaxed three-quarter pose against a plain warm gray studio backdrop. She wears a long camel wool coat open over a cream knit sweater, straight-leg dark indigo jeans, and white leather sneakers. Hair in a loose low bun, hands in the coat pockets. Soft even studio lighting from the front left, subtle floor shadow. Photoreal fashion catalog photography, full length, centered, no text.

An e-commerce catalog photo of a woman in her late twenties standing in a relaxed three-quarter pose against a plain warm gray studio backdrop. She wears a long camel wool coat open over a cream knit sweater, straight-leg dark indigo jeans, and white leather sneakers. Hair in a loose low bun, hands in the coat pockets. Soft even studio lighting from the front left, subtle floor shadow. Photoreal fashion catalog photography, full length, centered, no text.

An e-commerce catalog photo of a woman in her late twenties standing in a relaxed three-quarter pose against a plain warm gray studio backdrop. She wears a long camel wool coat open over a cream knit sweater, straight-leg dark indigo jeans, and white leather sneakers. Hair in a loose low bun, hands in the coat pockets. Soft even studio lighting from the front left, subtle floor shadow. Photoreal fashion catalog photography, full length, centered, no text.
The first two calls returned the same file, byte for byte. The third keeps every item the prompt named and changes what it left open, such as the cut of the coat and where she looks. A seed is how you keep a render you liked while you adjust the wording around it, and how you hand a colleague the exact image instead of a description of it.
Reproducibility needs promptExtend set to false. With extension on, the same seed and prompt return a different image on every call, because the rewrite that runs first is not seeded. Prompt extension shows the difference.
Tips
-
Set both dimensions on every request. Leaving
widthandheightout returns 512 × 512, a fraction of what the budget allows. -
Turn extension off while you iterate. With
promptExtendon you are tuning a prompt the model never sees, and two runs of the same wording will not match. -
Pick the shape before the wording. A strip and a portrait frame need different composition language, and the words only pay off if the canvas has room for them.
-
Quote every string that has to be spelled exactly. Write it in the case you want printed, and describe its size and position alongside it.
-
Keep a seed with every render you might need again. It costs nothing to send and it is the only way back to a specific image.