Nano Banana 2.1

Nano Banana 2.1 is Google's updated image generation and editing model in the Nano Banana 2 family. It improves graphic composition and subject consistency, and it follows complex prompts and edit instructions more accurately. It works from text alone or from as many as fourteen reference images plus one reference video, and it can ground a generation in live web and image search so the result reflects current facts and real visual references. It generates at 1K, 2K and 4K, with three levels of thinking that trade speed for reasoning depth, which suits layout-heavy design work, recurring characters and products, and precise multi-step edits.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
← All GuidesPrompting
How to write prompts for Nano Banana 2.1: briefs with counts and positions, layered descriptions, a system prompt for a house style, thinking levels, and sizes.
Introduction
Nano Banana 2.1 treats a prompt as a brief to carry out. It tracks how many of each object you named and where each one sits, then builds the frame to match. A short prompt still returns a finished image, with every unstated decision made by the model. A prompt that states the frame gets that frame.

A campaign photograph for a trail running brand. A woman in her thirties runs along a narrow ridge trail at sunrise, mid-stride with her left foot forward, wearing a coral windbreaker, black shorts, and a light gray hydration vest. The low sun sits behind her right shoulder and rims her hair and jacket with warm light. Valley fog fills the background below the ridge, and dry golden grass blurs in the foreground. Shot from trail level with a 35mm lens, shallow depth of field, photoreal, no text, no logos.
That prompt reads like a shot list: who, doing what, wearing what, where the light comes from, where the camera sits. This guide covers the request, counts and positions, layering a longer prompt, exclusions, a system prompt for a house style, the three thinking levels, and how to choose a size.
The request
Text-to-image needs a prompt and a size:
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'google:nano-banana@2.1',
positivePrompt: 'A campaign photograph for a trail running brand. A woman in her thirties runs along a narrow ridge trail at sunrise, mid-stride with her left foot forward, wearing a coral windbreaker, black shorts, and a light gray hydration vest. The low sun sits behind her right shoulder and rims her hair and jacket with warm light. Valley fog fills the background below the ridge, and dry golden grass blurs in the foreground. Shot from trail level with a 35mm lens, shallow depth of field, photoreal, no text, no logos.',
width: 2528,
height: 1696
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "google:nano-banana@2.1",
"positivePrompt": "A campaign photograph for a trail running brand. A woman in her thirties runs along a narrow ridge trail at sunrise, mid-stride with her left foot forward, wearing a coral windbreaker, black shorts, and a light gray hydration vest. The low sun sits behind her right shoulder and rims her hair and jacket with warm light. Valley fog fills the background below the ridge, and dry golden grass blurs in the foreground. Shot from trail level with a 35mm lens, shallow depth of field, photoreal, no text, no logos.",
"width": 2528,
"height": 1696
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "imageInference",
"taskUUID": "0b6a1f3e-7d52-4c1a-9e0f-3a8c5d2b7e41",
"model": "google:nano-banana@2.1",
"positivePrompt": "A campaign photograph for a trail running brand. A woman in her thirties runs along a narrow ridge trail at sunrise, mid-stride with her left foot forward, wearing a coral windbreaker, black shorts, and a light gray hydration vest. The low sun sits behind her right shoulder and rims her hair and jacket with warm light. Valley fog fills the background below the ridge, and dry golden grass blurs in the foreground. Shot from trail level with a 35mm lens, shallow depth of field, photoreal, no text, no logos.",
"width": 2528,
"height": 1696
}
]'runware run google:nano-banana@2.1 \
positivePrompt="A campaign photograph for a trail running brand. A woman in her thirties runs along a narrow ridge trail at sunrise, mid-stride with her left foot forward, wearing a coral windbreaker, black shorts, and a light gray hydration vest. The low sun sits behind her right shoulder and rims her hair and jacket with warm light. Valley fog fills the background below the ridge, and dry golden grass blurs in the foreground. Shot from trail level with a 35mm lens, shallow depth of field, photoreal, no text, no logos." \
width=2528 \
height=1696{
"taskType": "imageInference",
"taskUUID": "0b6a1f3e-7d52-4c1a-9e0f-3a8c5d2b7e41",
"model": "google:nano-banana@2.1",
"positivePrompt": "A campaign photograph for a trail running brand. A woman in her thirties runs along a narrow ridge trail at sunrise, mid-stride with her left foot forward, wearing a coral windbreaker, black shorts, and a light gray hydration vest. The low sun sits behind her right shoulder and rims her hair and jacket with warm light. Valley fog fills the background below the ridge, and dry golden grass blurs in the foreground. Shot from trail level with a 35mm lens, shallow depth of field, photoreal, no text, no logos.",
"width": 2528,
"height": 1696
}Response
[
{
"taskType": "imageInference",
"taskUUID": "0b6a1f3e-7d52-4c1a-9e0f-3a8c5d2b7e41",
"imageUUID": "c41d9a70-52be-4f08-a3c6-1e7f0b94d2a5",
"imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/c41d9a70-52be-4f08-a3c6-1e7f0b94d2a5.jpg"
}
]positivePrompt runs up to 45,000 characters, so a full creative brief fits without trimming. width and height come as a pair from a fixed list, covered in Choosing a size. Attaching reference images turns the same request into an edit or a composition, which the editing and multi-reference guides cover.
Counts and positions
A list of objects leaves the arrangement to the model. A number and a place for each object turns the list into a layout. Write the count next to the noun, and anchor each item to a region of the frame: left, center, top right, along the bottom edge.

An overhead photograph for a food delivery app, shot straight down onto a pale oak table. Exactly three bowls sit in a row across the middle of the frame: on the left a salmon poke bowl with avocado and edamame, in the center a chicken katsu curry with white rice, on the right a vegetable ramen with a halved soft-boiled egg. Exactly four wooden chopsticks lie in two pairs in the bottom right corner. A small white dish of soy sauce sits at the top left, and a folded sage green linen napkin at the top right. Soft daylight from the left, crisp shadows, photoreal food photography, no text.
The pattern is count, noun, region, repeated once per item: "exactly three bowls sit in a row across the middle", "exactly four wooden chopsticks lie in two pairs in the bottom right corner". The word "exactly" marks a count as a requirement. For a layout with many more items than this one, raise the thinking level.
Layering a longer prompt
A long prompt holds together when it moves in one direction, from the subject outward. Start with who or what the image is about, then what surrounds it, then how it is lit and photographed. The headshot prompt below has six layers:

A corporate headshot of a woman in her forties with short silver hair and a relaxed half smile, wearing a navy blazer over a white crew-neck top, standing in a bright open-plan office with glass partitions blurred behind her, framed from the chest up, slightly left of center, looking at the camera, soft window light from the right with a gentle shadow on the left cheek, 85mm lens, shallow depth of field, photoreal, natural skin texture.
Camera terms are worth their words. A focal length and a depth-of-field cue set the framing and the background blur more reliably than "close-up" or "professional". The layers are a checklist of what you could specify, and most prompts need only some of them. A product shot skips wardrobe, and a flat illustration skips the lens.
Leaving things out
This model has no negativePrompt field. An exclusion goes inside the prompt as a closing clause that starts with "Negative prompt:" and lists what to keep out. Here is one travel brochure shot, with and without the clause:

A travel brochure photograph of the Charles Bridge in Prague on a summer afternoon, looking along the bridge toward the Old Town bridge tower, warm light on the stone statues.

A travel brochure photograph of the Charles Bridge in Prague on a summer afternoon, looking along the bridge toward the Old Town bridge tower, warm light on the stone statues. Negative prompt: people, crowds, tourists
The two prompts differ only by the trailing clause. List nouns, not sentences: "people, crowds, tourists" reads as a set of things to exclude, the same way a dedicated field would take it.
A house style in the system prompt
settings.systemPrompt carries an instruction that applies on top of the prompt. It is the place for the part of a brief that never changes between requests: an illustration style, a palette, a background rule, a ban on text. With that in place, each prompt only has to name its subject.
You illustrate for a personal finance app. Every image is a flat vector spot illustration with thick rounded dark navy outlines and flat fills limited to coral, mustard yellow, teal, and off-white. No gradients, no shading, no text. One centered subject on a plain off-white background with a wide empty margin on every side.
A piggy bank with a single coin dropping into the slot.

A paper receipt under a magnifying glass.

A wall calendar with a ringing bell on one corner.
The three prompts are one sentence each, and none of them mentions a style or a palette:
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'google:nano-banana@2.1',
positivePrompt: 'A piggy bank with a single coin dropping into the slot.',
width: 2048,
height: 2048,
settings: {
systemPrompt: 'You illustrate for a personal finance app. Every image is a flat vector spot illustration with thick rounded dark navy outlines and flat fills limited to coral, mustard yellow, teal, and off-white. No gradients, no shading, no text. One centered subject on a plain off-white background with a wide empty margin on every side.'
}
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "google:nano-banana@2.1",
"positivePrompt": "A piggy bank with a single coin dropping into the slot.",
"width": 2048,
"height": 2048,
"settings": {
"systemPrompt": "You illustrate for a personal finance app. Every image is a flat vector spot illustration with thick rounded dark navy outlines and flat fills limited to coral, mustard yellow, teal, and off-white. No gradients, no shading, no text. One centered subject on a plain off-white background with a wide empty margin on every side."
}
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "imageInference",
"taskUUID": "5e2c8d14-a9f3-4b67-8c21-d0e47f6a3b98",
"model": "google:nano-banana@2.1",
"positivePrompt": "A piggy bank with a single coin dropping into the slot.",
"width": 2048,
"height": 2048,
"settings": {
"systemPrompt": "You illustrate for a personal finance app. Every image is a flat vector spot illustration with thick rounded dark navy outlines and flat fills limited to coral, mustard yellow, teal, and off-white. No gradients, no shading, no text. One centered subject on a plain off-white background with a wide empty margin on every side."
}
}
]'runware run google:nano-banana@2.1 \
positivePrompt="A piggy bank with a single coin dropping into the slot." \
width=2048 \
height=2048 \
settings.systemPrompt="You illustrate for a personal finance app. Every image is a flat vector spot illustration with thick rounded dark navy outlines and flat fills limited to coral, mustard yellow, teal, and off-white. No gradients, no shading, no text. One centered subject on a plain off-white background with a wide empty margin on every side."{
"taskType": "imageInference",
"taskUUID": "5e2c8d14-a9f3-4b67-8c21-d0e47f6a3b98",
"model": "google:nano-banana@2.1",
"positivePrompt": "A piggy bank with a single coin dropping into the slot.",
"width": 2048,
"height": 2048,
"settings": {
"systemPrompt": "You illustrate for a personal finance app. Every image is a flat vector spot illustration with thick rounded dark navy outlines and flat fills limited to coral, mustard yellow, teal, and off-white. No gradients, no shading, no text. One centered subject on a plain off-white background with a wide empty margin on every side."
}
}Keep the system prompt to rules that hold for every image, and leave anything that varies to the prompt. A subject described in the system prompt shows up in every image of the batch.
Thinking levels
settings.thinkingLevel sets how long the model reasons about the brief before it renders. It takes minimal, medium, or high, and medium is the default. Reasoning is where the model reconciles counts and positions against each other, so the level matters most when a prompt stacks many constraints and least for a single subject on a plain background.
The prompt below is a kit-contents shot for a hardware store listing: seven kinds of tool, each with a count, in a fixed grid. It ran once at each level.

A product photograph for an online hardware store showing the full contents of a home tool kit, laid out in a neat grid on a dark gray felt mat and shot from directly above. Top row, left to right: exactly four screwdrivers with orange handles in descending size, then one claw hammer with a wooden handle. Middle row, left to right: exactly two pairs of pliers with red grips, one yellow tape measure, and one small torpedo level. Bottom row: exactly six hex keys fanned out in a quarter circle on the left, and one utility knife on the right. Even soft studio lighting, every tool fully inside the frame with space between items, photoreal, no text, no logos.

A product photograph for an online hardware store showing the full contents of a home tool kit, laid out in a neat grid on a dark gray felt mat and shot from directly above. Top row, left to right: exactly four screwdrivers with orange handles in descending size, then one claw hammer with a wooden handle. Middle row, left to right: exactly two pairs of pliers with red grips, one yellow tape measure, and one small torpedo level. Bottom row: exactly six hex keys fanned out in a quarter circle on the left, and one utility knife on the right. Even soft studio lighting, every tool fully inside the frame with space between items, photoreal, no text, no logos.

A product photograph for an online hardware store showing the full contents of a home tool kit, laid out in a neat grid on a dark gray felt mat and shot from directly above. Top row, left to right: exactly four screwdrivers with orange handles in descending size, then one claw hammer with a wooden handle. Middle row, left to right: exactly two pairs of pliers with red grips, one yellow tape measure, and one small torpedo level. Bottom row: exactly six hex keys fanned out in a quarter circle on the left, and one utility knife on the right. Even soft studio lighting, every tool fully inside the frame with space between items, photoreal, no text, no logos.
Count the hex keys in each version. minimal returned eight, in two groups of four. medium and high both returned the six the prompt asked for, and only high laid the tools out in the three rows as written.
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'google:nano-banana@2.1',
positivePrompt: 'A product photograph for an online hardware store showing the full contents of a home tool kit, laid out in a neat grid on a dark gray felt mat and shot from directly above. Top row, left to right: exactly four screwdrivers with orange handles in descending size, then one claw hammer with a wooden handle. Middle row, left to right: exactly two pairs of pliers with red grips, one yellow tape measure, and one small torpedo level. Bottom row: exactly six hex keys fanned out in a quarter circle on the left, and one utility knife on the right. Even soft studio lighting, every tool fully inside the frame with space between items, photoreal, no text, no logos.',
width: 2528,
height: 1696,
settings: {
thinkingLevel: 'high'
}
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "google:nano-banana@2.1",
"positivePrompt": "A product photograph for an online hardware store showing the full contents of a home tool kit, laid out in a neat grid on a dark gray felt mat and shot from directly above. Top row, left to right: exactly four screwdrivers with orange handles in descending size, then one claw hammer with a wooden handle. Middle row, left to right: exactly two pairs of pliers with red grips, one yellow tape measure, and one small torpedo level. Bottom row: exactly six hex keys fanned out in a quarter circle on the left, and one utility knife on the right. Even soft studio lighting, every tool fully inside the frame with space between items, photoreal, no text, no logos.",
"width": 2528,
"height": 1696,
"settings": {
"thinkingLevel": "high"
}
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "imageInference",
"taskUUID": "9a7d3c52-1e84-4f0b-b6a9-27c5e8d1f430",
"model": "google:nano-banana@2.1",
"positivePrompt": "A product photograph for an online hardware store showing the full contents of a home tool kit, laid out in a neat grid on a dark gray felt mat and shot from directly above. Top row, left to right: exactly four screwdrivers with orange handles in descending size, then one claw hammer with a wooden handle. Middle row, left to right: exactly two pairs of pliers with red grips, one yellow tape measure, and one small torpedo level. Bottom row: exactly six hex keys fanned out in a quarter circle on the left, and one utility knife on the right. Even soft studio lighting, every tool fully inside the frame with space between items, photoreal, no text, no logos.",
"width": 2528,
"height": 1696,
"settings": {
"thinkingLevel": "high"
}
}
]'runware run google:nano-banana@2.1 \
positivePrompt="A product photograph for an online hardware store showing the full contents of a home tool kit, laid out in a neat grid on a dark gray felt mat and shot from directly above. Top row, left to right: exactly four screwdrivers with orange handles in descending size, then one claw hammer with a wooden handle. Middle row, left to right: exactly two pairs of pliers with red grips, one yellow tape measure, and one small torpedo level. Bottom row: exactly six hex keys fanned out in a quarter circle on the left, and one utility knife on the right. Even soft studio lighting, every tool fully inside the frame with space between items, photoreal, no text, no logos." \
width=2528 \
height=1696 \
settings.thinkingLevel=high{
"taskType": "imageInference",
"taskUUID": "9a7d3c52-1e84-4f0b-b6a9-27c5e8d1f430",
"model": "google:nano-banana@2.1",
"positivePrompt": "A product photograph for an online hardware store showing the full contents of a home tool kit, laid out in a neat grid on a dark gray felt mat and shot from directly above. Top row, left to right: exactly four screwdrivers with orange handles in descending size, then one claw hammer with a wooden handle. Middle row, left to right: exactly two pairs of pliers with red grips, one yellow tape measure, and one small torpedo level. Bottom row: exactly six hex keys fanned out in a quarter circle on the left, and one utility knife on the right. Even soft studio lighting, every tool fully inside the frame with space between items, photoreal, no text, no logos.",
"width": 2528,
"height": 1696,
"settings": {
"thinkingLevel": "high"
}
}minimal suits drafts and simple scenes, where it returns sooner. high suits dense layouts and long copy, and it takes the longest. Stay on the default until an output misses a count or misplaces an item, then rerun that prompt on high.
Choosing a size
width and height are a pair from a fixed list. The list is 14 aspect ratios, each in three tiers. The 1K pairs are below. The 2K tier doubles both sides and the 4K tier doubles them again, so 16:9 is 1376 × 768, 2752 × 1536, or 5504 × 3072.
| Aspect ratio | Landscape | Portrait |
|---|---|---|
| 1:1 | 1024 × 1024 | |
| 5:4 and 4:5 | 1152 × 928 | 928 × 1152 |
| 4:3 and 3:4 | 1200 × 896 | 896 × 1200 |
| 3:2 and 2:3 | 1264 × 848 | 848 × 1264 |
| 16:9 and 9:16 | 1376 × 768 | 768 × 1376 |
| 21:9 | 1584 × 672 | |
| 4:1 and 1:4 | 2064 × 512 | 512 × 2064 |
| 8:1 and 1:8 | 2928 × 352 | 352 × 2928 |
The 4:1 strip covers formats that usually mean cropping a wider render, such as site headers and email banners. Describe it as one scene with a centerpiece and something on either side, so the subject spreads across the full width:

A website header banner for an outdoor furniture store's summer sale: a single wide photograph of a sunlit stone terrace above a calm blue sea. A round teak dining table with four chairs under a large off-white parasol stands in the middle of the terrace. To its left, a teak lounge chair with a cream cushion and a striped throw. To its right, a pair of terracotta planters with olive trees. A low stone wall runs along the back of the whole terrace, in front of one unbroken horizon. Bright midday light, photoreal, no text, no people.
That banner is the 2K 4:1 pair, 4128 × 1024. Pick the tier by where the image ships: 1K for drafts and thumbnails, 2K for the web, 4K for print or for a frame you plan to crop.
resolution is the alternative to a pair, and it only applies when a reference image is attached. It sets the tier and takes the aspect ratio from the reference, which is how the editing guide sizes its results.
Tips
-
Put a count and a region next to every item. "Exactly three bowls in a row across the middle" is a layout. "Some bowls on a table" is a suggestion.
-
Move from the subject outward. Subject, wardrobe, setting, framing, lighting, camera. A prompt in that order drops fewer details than the same facts in a run-on sentence.
-
Write exclusions as a closing clause. End the prompt with "Negative prompt:" and a list of nouns.
-
Move the constant part of a brief into the system prompt. Style and palette rules go in
settings.systemPrompt, and each prompt names only its subject. -
Raise the thinking level when an output misses a count. The default handles most prompts. Dense layouts and long copy are the case for
high. -
Pick the pair, then the tier. The aspect ratio follows the placement, and the tier follows the delivery size.