live
MODEL IDalibaba:qwen-image@2.1-pro

Qwen-Image-2.1-Pro

Alibaba
by

Qwen-Image-2.1-Pro is the Pro tier of Alibaba's Qwen-Image-2.1 family, a unified model for text-to-image generation and prompt-guided image editing. It composes and edits from up to 10 reference images and can rewrite prompts with an LLM before generation, in a direct or agent mode with optional thinking. It outputs from 0.26 to 4.19 megapixels at aspect ratios from 1:8 to 8:1, suited to product imagery, marketing visuals, and multi-image compositions.

Qwen-Image-2.1-Pro

Prompt extension

How promptExtend works on Qwen-Image-2.1-Pro: what the LLM rewrite adds, how direct and agent differ, where thinking fits, and why an extended prompt breaks seeds.

Introduction

Qwen-Image-2.1-Pro runs your prompt through an LLM before it generates, expanding a short description into a full one. The rewrite is on unless you turn it off, so the prompt in your request is usually not the prompt the model rendered.

The hero below came from six words, a rooftop yoga class at sunrise. Everything else in it was written by the rewrite:

The water-tower skyline, the size of the class, the yoga blocks, and the bottles beside the mats were never in the request. You get a finished scene from a one-line brief, and every choice in it belongs to the rewrite.

Request shape

Three settings control the rewrite. promptExtend defaults to on, promptExtendMode picks between direct and agent, and thinking lets the rewriting LLM reason before it writes:

Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'alibaba:qwen-image@2.1-pro',
  positivePrompt: 'a rooftop yoga class at sunrise',
  width: 2496,
  height: 1664,
  settings: {
    promptExtend: true,
    promptExtendMode: 'agent',
    thinking: true
  }
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "alibaba:qwen-image@2.1-pro",
            "positivePrompt": "a rooftop yoga class at sunrise",
            "width": 2496,
            "height": 1664,
            "settings": {
                "promptExtend": True,
                "promptExtendMode": "agent",
                "thinking": True
            }
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "imageInference",
      "taskUUID": "d9e826f2-de4f-4183-a787-eaeaa8169192",
      "model": "alibaba:qwen-image@2.1-pro",
      "positivePrompt": "a rooftop yoga class at sunrise",
      "width": 2496,
      "height": 1664,
      "settings": {
        "promptExtend": true,
        "promptExtendMode": "agent",
        "thinking": true
      }
    }
  ]'
runware run alibaba:qwen-image@2.1-pro \
  positivePrompt="a rooftop yoga class at sunrise" \
  width=2496 \
  height=1664 \
  settings.promptExtend=true \
  settings.promptExtendMode=agent \
  settings.thinking=true
{
  "taskType": "imageInference",
  "taskUUID": "d9e826f2-de4f-4183-a787-eaeaa8169192",
  "model": "alibaba:qwen-image@2.1-pro",
  "positivePrompt": "a rooftop yoga class at sunrise",
  "width": 2496,
  "height": 1664,
  "settings": {
    "promptExtend": true,
    "promptExtendMode": "agent",
    "thinking": true
  }
}
Response
[
  {
    "taskType": "imageInference",
    "taskUUID": "d9e826f2-de4f-4183-a787-eaeaa8169192",
    "imageUUID": "8edd394e-f696-4482-948c-6b1968623711",
    "imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/8edd394e-f696-4482-948c-6b1968623711.jpg"
  }
]

Two combinations are rejected. thinking belongs to the rewrite, so sending it with promptExtend: false fails. And agent is text-to-image only: a request that carries inputs.referenceImages must use direct.

What the rewrite adds

Both images below came from a hotel lobby. The only difference is whether the rewrite ran:

Both are lobbies. With the rewrite on you get a scene the rewrite decided you meant, with its own materials, lighting, furniture, and time of day.

The rewrite adds subjects, and added subjects can carry text and branding nobody requested. The extended render has a hotel name and a crest on the wall behind the desk. Check for that before an extended image goes anywhere commercial.

direct and agent

promptExtendMode picks the rewriting strategy: direct, the default, or agent. Both images came from a launch ad for wireless headphones:

Both modes turned six words into a finished ad, and every word in both ads was invented by the rewrite: a brand, a headline, feature claims, a pre-order button. The agent render goes further into specifics, with a model name, a $399 price, a ship date, and certification badges.

Treat an extended ad as a layout proposal. When the image is meant to ship, put the real name and price in the prompt, in quotes, so the rewrite has nothing left to make up.

Thinking

This setting gives the rewriting LLM a reasoning pass before it writes the expanded prompt, at the cost of longer generation time. Both images below used agent on an infographic showing the five steps to repot a houseplant:

Both renders deliver five numbered steps in a workable order, plus side panels neither request asked for. On a brief like this one, agent already plans the content without the thinking pass, so the extra time buys little. Compare a pair on your own prompts before paying for it on every call.

Extension on an edit

The rewrite also runs on edits. The source below is a summer listing photo, and both edits used the instruction make it winter:

Both edits changed the season and kept the house. With extension on, the snow also buries the path and the road: the rewrite decided how much winter you meant. On an edit the source image holds the subject in place, so what the rewrite changes is degree. State the degree in the instruction when it matters, or turn extension off.

Extension and reproducibility

With extension off, a fixed seed returns the same image on every call, as Prompting shows. Turn extension on, keep the seed and the prompt identical, and two calls diverge:

Same brief, two different people. The rewrite is not covered by your seed, so it writes a fresh prompt on every call, and the seed then locks the noise for a description that keeps moving. For a team page that needs a retake of one person, that makes the seed useless.

"settings": {
  "promptExtend": false
}

That block is the whole fix. Anything that has to be repeatable, whether a regression test or a client revision, needs extension off.

Tips

  1. Leave it on for exploration. A short prompt with agent is a fast way to find a direction you would not have written.

  2. Turn it off before you commit. Once the composition matters, extension is a second author working from a copy of your brief.

  3. Test the thinking pass before you rely on it. On a structured brief, agent produced the same five-step layout with and without it, and the pass adds generation time.

  4. Check extended outputs for invented copy. The rewrite adds brand names and prices on its own, which is a problem for anything meant to ship.

  5. Keep the default mode on anything with references. agent is rejected as soon as the request carries inputs.referenceImages.