live
MODEL IDgoogle:nano-banana@2.1

Nano Banana 2.1

Google
by

Nano Banana 2.1 is Google's updated image generation and editing model in the Nano Banana 2 family. It improves graphic composition and subject consistency, and it follows complex prompts and edit instructions more accurately. It works from text alone or from as many as fourteen reference images plus one reference video, and it can ground a generation in live web and image search so the result reflects current facts and real visual references. It generates at 1K, 2K and 4K, with three levels of thinking that trade speed for reasoning depth, which suits layout-heavy design work, recurring characters and products, and precise multi-step edits.

Nano Banana 2.1

Composing from multiple reference images

How to combine up to 14 reference images with Nano Banana 2.1: giving each image a role, following a layout sketch, borrowing a style, and reading a video.

Introduction

Putting a person and a product into a location neither was shot in is normally a compositing job. You cut out each element, place it, then match scale and light by hand, and every revision repeats the work.

Nano Banana 2.1 does it in one request. Each element goes into inputs.referenceImages as its own image, the prompt says which image plays which part, and the model returns one frame with scale and lighting already reconciled. A request takes up to 14 images.

Built from 4 references
  • Image 1: rider
  • Image 2: bike
  • Image 3: helmet
  • Image 4: location

The rider never sat on that bike, and nobody was on the path. For a bike brand, that is a campaign image built from catalog assets.

This guide covers the request, how to give each image a role, arranging a set of products, following a layout sketch, borrowing a style, and using a video as the reference. Holding one subject across a series of images is the reverse job, covered in subject consistency.

The request

References go into inputs.referenceImages as URLs, UUIDs, data URIs, or base64 strings:

Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'google:nano-banana@2.1',
  positivePrompt: 'The man from image 1 riding the green e-bike from image 2 along the waterfront cycle path from image 4, wearing the cream helmet from image 3 with the straps fastened. He rides toward the camera at a slight angle, both hands on the grips, with the harbor on the left of the frame. Match the golden hour light and long shadows of image 4 on the rider and the bike. Keep his face, beard, and clothing from image 1, and keep the bike\'s frame color, saddle, basket, and tires from image 2. Photoreal lifestyle campaign photography, shallow depth of field.',
  width: 2528,
  height: 1696,
  inputs: {
    referenceImages: [
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/f4a19c63-0e7b-4d52-8b3a-c61d5e9f2a07.jpg',
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/1c7e0b48-a5d3-4f69-9e12-3b8a6d4c0f95.jpg',
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/95d2f6a1-3b84-4c0e-a7f5-0e9c1b7d6a34.jpg',
      'https://im.runware.ai/image/os/a14d18/ws/2/ii/60b3a8e7-f1c9-4d25-b4a6-8d2e5f0c9b71.jpg'
    ]
  }
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "google:nano-banana@2.1",
            "positivePrompt": "The man from image 1 riding the green e-bike from image 2 along the waterfront cycle path from image 4, wearing the cream helmet from image 3 with the straps fastened. He rides toward the camera at a slight angle, both hands on the grips, with the harbor on the left of the frame. Match the golden hour light and long shadows of image 4 on the rider and the bike. Keep his face, beard, and clothing from image 1, and keep the bike's frame color, saddle, basket, and tires from image 2. Photoreal lifestyle campaign photography, shallow depth of field.",
            "width": 2528,
            "height": 1696,
            "inputs": {
                "referenceImages": [
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/f4a19c63-0e7b-4d52-8b3a-c61d5e9f2a07.jpg",
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/1c7e0b48-a5d3-4f69-9e12-3b8a6d4c0f95.jpg",
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/95d2f6a1-3b84-4c0e-a7f5-0e9c1b7d6a34.jpg",
                    "https://im.runware.ai/image/os/a14d18/ws/2/ii/60b3a8e7-f1c9-4d25-b4a6-8d2e5f0c9b71.jpg"
                ]
            }
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "imageInference",
      "taskUUID": "b82e5d0a-4c17-4f93-a6d1-7e3f9c0b5a28",
      "model": "google:nano-banana@2.1",
      "positivePrompt": "The man from image 1 riding the green e-bike from image 2 along the waterfront cycle path from image 4, wearing the cream helmet from image 3 with the straps fastened. He rides toward the camera at a slight angle, both hands on the grips, with the harbor on the left of the frame. Match the golden hour light and long shadows of image 4 on the rider and the bike. Keep his face, beard, and clothing from image 1, and keep the bike's frame color, saddle, basket, and tires from image 2. Photoreal lifestyle campaign photography, shallow depth of field.",
      "width": 2528,
      "height": 1696,
      "inputs": {
        "referenceImages": [
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/f4a19c63-0e7b-4d52-8b3a-c61d5e9f2a07.jpg",
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/1c7e0b48-a5d3-4f69-9e12-3b8a6d4c0f95.jpg",
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/95d2f6a1-3b84-4c0e-a7f5-0e9c1b7d6a34.jpg",
          "https://im.runware.ai/image/os/a14d18/ws/2/ii/60b3a8e7-f1c9-4d25-b4a6-8d2e5f0c9b71.jpg"
        ]
      }
    }
  ]'
runware run google:nano-banana@2.1 \
  positivePrompt="The man from image 1 riding the green e-bike from image 2 along the waterfront cycle path from image 4, wearing the cream helmet from image 3 with the straps fastened. He rides toward the camera at a slight angle, both hands on the grips, with the harbor on the left of the frame. Match the golden hour light and long shadows of image 4 on the rider and the bike. Keep his face, beard, and clothing from image 1, and keep the bike's frame color, saddle, basket, and tires from image 2. Photoreal lifestyle campaign photography, shallow depth of field." \
  width=2528 \
  height=1696 \
  inputs.referenceImages.0=https://im.runware.ai/image/os/a14d18/ws/2/ii/f4a19c63-0e7b-4d52-8b3a-c61d5e9f2a07.jpg \
  inputs.referenceImages.1=https://im.runware.ai/image/os/a14d18/ws/2/ii/1c7e0b48-a5d3-4f69-9e12-3b8a6d4c0f95.jpg \
  inputs.referenceImages.2=https://im.runware.ai/image/os/a14d18/ws/2/ii/95d2f6a1-3b84-4c0e-a7f5-0e9c1b7d6a34.jpg \
  inputs.referenceImages.3=https://im.runware.ai/image/os/a14d18/ws/2/ii/60b3a8e7-f1c9-4d25-b4a6-8d2e5f0c9b71.jpg
{
  "taskType": "imageInference",
  "taskUUID": "b82e5d0a-4c17-4f93-a6d1-7e3f9c0b5a28",
  "model": "google:nano-banana@2.1",
  "positivePrompt": "The man from image 1 riding the green e-bike from image 2 along the waterfront cycle path from image 4, wearing the cream helmet from image 3 with the straps fastened. He rides toward the camera at a slight angle, both hands on the grips, with the harbor on the left of the frame. Match the golden hour light and long shadows of image 4 on the rider and the bike. Keep his face, beard, and clothing from image 1, and keep the bike's frame color, saddle, basket, and tires from image 2. Photoreal lifestyle campaign photography, shallow depth of field.",
  "width": 2528,
  "height": 1696,
  "inputs": {
    "referenceImages": [
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/f4a19c63-0e7b-4d52-8b3a-c61d5e9f2a07.jpg",
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/1c7e0b48-a5d3-4f69-9e12-3b8a6d4c0f95.jpg",
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/95d2f6a1-3b84-4c0e-a7f5-0e9c1b7d6a34.jpg",
      "https://im.runware.ai/image/os/a14d18/ws/2/ii/60b3a8e7-f1c9-4d25-b4a6-8d2e5f0c9b71.jpg"
    ]
  }
}
Response
[
  {
    "taskType": "imageInference",
    "taskUUID": "b82e5d0a-4c17-4f93-a6d1-7e3f9c0b5a28",
    "imageUUID": "2a6f9d31-8c05-4e7b-b1d4-5f0a3c8e6b92",
    "imageURL": "https://im.runware.ai/image/os/a14d18/ws/2/ii/2a6f9d31-8c05-4e7b-b1d4-5f0a3c8e6b92.jpg"
  }
]

The array order is the numbering. "Image 1" in the prompt is the first entry of the array, and "image 2" is the second. width and height set the output frame on their own, whatever the shapes of the references.

Giving each image a role

A reference can play one of four parts: a subject to keep, a product to place, a setting to build in, or a look to copy. The prompt has to say which, because one photo of a room could be a location or a color palette. Four habits keep the roles unambiguous:

  • Name each image by position and by content. "The green e-bike from image 2" survives a reordered array better than "image 2" alone, and it tells the model what to look for in that image.
  • Say what to take from it. "Keep the bike's frame color, saddle, basket, and tires" lists the parts of the reference that have to arrive intact.
  • Say how the pieces relate. Riding, wearing, standing on, resting beside. The references cannot express a relationship, so the prompt has to.
  • Describe the light once, for the final scene. The references can come from different shoots. One lighting sentence, anchored to the location image, covers every element.

A product set in one frame

The same pattern scales to a full range. Six packshots go in, and one arranged flat-lay comes out:

Built from 6 references
  • Image 1
  • Image 2
  • Image 3
  • Image 4
  • Image 5
  • Image 6

A flat-lay has room to fill, and without a count a product can show up twice. Give every reference a position, and state that each product appears exactly once.

Following a layout sketch

A reference can also be a layout. Draw the composition as rough boxes, pass the drawing along with the assets, and tell the model to build the finished design on that plan:

Built from 3 references
  • Image 1: sketch
  • Image 2: product
  • Image 3: logo

The sketch carries no style at all, only positions. Words written in the sketch become handles: the prompt points at "where the sketch says PRODUCT" and never describes a coordinate. A designer's thumbnail and a folder of assets are enough for a first comp.

Borrowing a style

Give one image the role of content and another the role of style, and the model redraws the first in the manner of the second:

Built from 2 references
  • Image 1: content
  • Image 2: style

The closing clause, "take nothing else from it", keeps the lighthouse out of the result. A style reference is still a picture of something, and a prompt that does not scope it can carry part of that something across.

A video as the reference

inputs.referenceVideos takes up to 10 videos, each as a URL, a UUID, or a public YouTube link. The model watches a clip and works from what happens in it. The request below passes one, which turns a video into the brief for its own thumbnail:

Built from 1 reference
  • Video

    A cooking video shot from a fixed overhead angle on a white marble counter: two hands toss fresh fusilli with bright green pesto in a wide steel pan, then tip the pasta into a white bowl and scatter pine nuts and a few basil leaves on top. Bright even kitchen light, photoreal. Sound: a soft sizzle and the scrape of a wooden spoon.

Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'google:nano-banana@2.1',
  positivePrompt: 'Design a video thumbnail for the cooking video. Show a close three-quarter view of the finished bowl of pesto pasta from the end of the video, with the same bowl, pasta shape, and toppings, on the same marble counter. On the left third of the frame, set the title "15-MINUTE PESTO PASTA" in heavy white capitals with a thin dark outline, stacked on three lines. In the top right corner, a round yellow badge with "EASY" in dark capitals. Bright, saturated, appetizing food photography.',
  width: 2752,
  height: 1536,
  inputs: {
    referenceVideos: [
      'https://vm.runware.ai/video/os/a14d18/ws/2/vi/7d3c9a15-e2b6-4f80-a4c7-9b1e0f5d3a28.mp4'
    ]
  }
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "google:nano-banana@2.1",
            "positivePrompt": "Design a video thumbnail for the cooking video. Show a close three-quarter view of the finished bowl of pesto pasta from the end of the video, with the same bowl, pasta shape, and toppings, on the same marble counter. On the left third of the frame, set the title \"15-MINUTE PESTO PASTA\" in heavy white capitals with a thin dark outline, stacked on three lines. In the top right corner, a round yellow badge with \"EASY\" in dark capitals. Bright, saturated, appetizing food photography.",
            "width": 2752,
            "height": 1536,
            "inputs": {
                "referenceVideos": [
                    "https://vm.runware.ai/video/os/a14d18/ws/2/vi/7d3c9a15-e2b6-4f80-a4c7-9b1e0f5d3a28.mp4"
                ]
            }
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "imageInference",
      "taskUUID": "c05f7e2b-91a4-4d38-b6c0-4a8e1d3f9b67",
      "model": "google:nano-banana@2.1",
      "positivePrompt": "Design a video thumbnail for the cooking video. Show a close three-quarter view of the finished bowl of pesto pasta from the end of the video, with the same bowl, pasta shape, and toppings, on the same marble counter. On the left third of the frame, set the title \"15-MINUTE PESTO PASTA\" in heavy white capitals with a thin dark outline, stacked on three lines. In the top right corner, a round yellow badge with \"EASY\" in dark capitals. Bright, saturated, appetizing food photography.",
      "width": 2752,
      "height": 1536,
      "inputs": {
        "referenceVideos": [
          "https://vm.runware.ai/video/os/a14d18/ws/2/vi/7d3c9a15-e2b6-4f80-a4c7-9b1e0f5d3a28.mp4"
        ]
      }
    }
  ]'
runware run google:nano-banana@2.1 \
  positivePrompt="Design a video thumbnail for the cooking video. Show a close three-quarter view of the finished bowl of pesto pasta from the end of the video, with the same bowl, pasta shape, and toppings, on the same marble counter. On the left third of the frame, set the title \"15-MINUTE PESTO PASTA\" in heavy white capitals with a thin dark outline, stacked on three lines. In the top right corner, a round yellow badge with \"EASY\" in dark capitals. Bright, saturated, appetizing food photography." \
  width=2752 \
  height=1536 \
  inputs.referenceVideos.0=https://vm.runware.ai/video/os/a14d18/ws/2/vi/7d3c9a15-e2b6-4f80-a4c7-9b1e0f5d3a28.mp4
{
  "taskType": "imageInference",
  "taskUUID": "c05f7e2b-91a4-4d38-b6c0-4a8e1d3f9b67",
  "model": "google:nano-banana@2.1",
  "positivePrompt": "Design a video thumbnail for the cooking video. Show a close three-quarter view of the finished bowl of pesto pasta from the end of the video, with the same bowl, pasta shape, and toppings, on the same marble counter. On the left third of the frame, set the title \"15-MINUTE PESTO PASTA\" in heavy white capitals with a thin dark outline, stacked on three lines. In the top right corner, a round yellow badge with \"EASY\" in dark capitals. Bright, saturated, appetizing food photography.",
  "width": 2752,
  "height": 1536,
  "inputs": {
    "referenceVideos": [
      "https://vm.runware.ai/video/os/a14d18/ws/2/vi/7d3c9a15-e2b6-4f80-a4c7-9b1e0f5d3a28.mp4"
    ]
  }
}

The clip was shot from directly overhead, and the thumbnail asks for a three-quarter view. The video supplies the content, and the prompt does the design work, from the camera angle to the title. The steel pan and the striped towel at the edges of the thumbnail come from the clip, and the prompt mentions neither. The same request shape turns a product demo into a store banner, or a tutorial into a step-by-step graphic.

Tips

  1. Number the images in the prompt. "The helmet from image 3" maps to the third entry of the array.

  2. Say what to take from each reference. List the parts that must arrive intact, and scope a style image with "take nothing else from it".

  3. Shoot references plain. A subject on a clean backdrop gives the model less to untangle than one already inside a busy scene.

  4. Describe the final light once. One lighting sentence for the output scene relights every element to match.

  5. Count the products in a set. A position for each reference and "exactly once" keep a flat-lay from repeating or dropping an item.

  6. Label the boxes in a layout sketch. The labels are what the prompt refers to.