Calling models with Python

Run inference from Python: transports, typed parameters, validation, concurrency, streaming, and the utilities around them.

Introduction

Everything below assumes a client from the overview, which is where installation and authentication live. What follows is how you drive it: which transport carries the request, how the parameters are typed, and what the SDK does with more than one call at a time.

Choosing a transport

The SDK ships two transports behind the same API. They differ in connection model and in how the server delivers the result.

TransportWhen to use it
restOne-off requests, serverless functions, short-lived processes. No persistent socket.
websocket (default)Many requests per process, lower per-call latency, push-based progress on long-running tasks.

Construct with the transport you want:

client = Runware(api_key="your-api-key", transport="rest")
# or
client = Runware(api_key="your-api-key", transport="websocket")

Either transport supports two delivery modes, set with deliveryMethod. Sync waits for the result and returns it in one round trip, which works well for fast tasks like image inference. Async returns a task UUID immediately and the SDK polls for completion, which is what you want for video and other long-running operations.

# Sync (recommended for image inference and other fast tasks)
images = await client.run({
    "taskType": "imageInference",
    "model": "runware:101@1",
    "positivePrompt": "A coastal town at dusk",
    "width": 1024,
    "height": 1024,
    "deliveryMethod": "sync",
})

# Async (the SDK polls for you; recommended for video)
videos = await client.run({
    "taskType": "videoInference",
    "model": "google:3@3",
    "positivePrompt": "Waves crashing on a beach",
    "width": 1280,
    "height": 720,
    "duration": 8,
    # deliveryMethod defaults to "async"
})

Over WebSocket, requests and results travel on the same persistent socket. With async delivery the SDK gets a taskUUID and polls for completion over that connection. With sync the result comes straight back. Either way, the SDK reconciles each frame with its awaiting call.

Typed parameters

When you know the architecture you're targeting, import the matching TypedDict and the SDK gives you compile-time validation in your editor:

from runware import Runware
from runware.types.task_map import SdxlArchParams

params: SdxlArchParams = {
    "model": "civitai:133005@782002",
    "taskType": "imageInference",
    "positivePrompt": "A professional headshot portrait",
    "negativePrompt": "blurry, distorted",
    "width": 1024,
    "height": 1024,
    "steps": 30,
}

async with Runware() as client:
    images = await client.run(params)

The task_map module ships one TypedDict per supported architecture and curated model. Pyright and mypy flag wrong field names and wrong value types before you run the code.

Schema validation

The SDK can validate your request against the model's JSON Schema before it leaves the process. Mistakes (a missing model, or an inputImage passed as bytes instead of a URL) surface as a typed exception with the offending field name, not as a 400 from the server hundreds of milliseconds later.

Validation is off by default. Turn it on for a call (or globally via validate=True in the Runware constructor) when you want the SDK to check a payload against the model's schema before sending:

from runware import Runware, RunOptions

async with Runware() as client:
    result = await client.run(
        payload,
        RunOptions(validate=True),
    )

Concurrent requests

The SDK is async all the way down. asyncio.gather is the canonical way to fan out:

import asyncio
from runware import Runware

async def main():
    async with Runware(transport="websocket") as client:
        await client.connect()

        results = await asyncio.gather(
            client.run({
                "taskType": "imageInference",
                "model": "runware:101@1",
                "positivePrompt": "Abstract digital art",
                "width": 1024,
                "height": 1024,
            }),
            client.run({
                "taskType": "imageInference",
                "model": "runware:101@1",
                "positivePrompt": "A neon-lit alley at night",
                "width": 1024,
                "height": 1024,
            }),
            client.run({
                "taskType": "imageBackgroundRemoval",
                "inputImage": "https://example.com/portrait.jpg",
            }),
        )

asyncio.run(main())

WebSocket has a real edge here. All three requests share one socket and the responses stream back independently. REST would issue three HTTP requests in parallel and tear them down after each.

LLM streaming

Text inference supports Server-Sent Events streaming for low-latency generation. Call client.stream() and iterate over the resulting TextStream:

from runware import Runware

async def main():
    async with Runware() as client:
        stream = await client.stream({
            "taskType": "textInference",
            "model": "minimax:m2.7@0",
            "messages": [
                {"role": "user", "content": "Write a haiku about the ocean."},
            ],
        })

        async for delta in stream.text_stream:
            print(delta, end="", flush=True)

        result = await stream.result()
        print(f"\nFinish reason: {result.finish_reason}")

import asyncio
asyncio.run(main())

The stream exposes two iterators (text_stream and reasoning_stream) plus a result() coroutine that yields the final accumulated text, finish reason, and usage stats. Iteration errors surface as typed exceptions, so a half-truncated stream never silently ends.

stream() handles a single completion only. Pass numberResults greater than 1 and it raises, use run() for batch text generation.

Content namespace

client.content reaches Runware's public model catalog and does not consume credits. Use it to list curated models, fetch pricing, pull example payloads, or browse the catalog by collection or creator.

async with Runware() as client:
    # Search the curated catalog
    models = await client.content.list_models({
        "category": "image",
        "creator": "black-forest-labs",
        "search": "flux dev",
    })

    # Inspect one
    flux = await client.content.get_model("flux-1-dev")
    print(flux["headline"])

    # Pull curated examples to seed prompts
    examples = await client.content.get_model_examples("flux-1-dev")

    # Pricing for budget-driven decisions
    pricing = await client.content.get_model_pricing("flux-1-dev")

    # Browse the capability taxonomy, collections, and creators
    capabilities = await client.content.list_capabilities()
    collections = await client.content.list_collections({"category": "image"})
    creators = await client.content.list_creators()

The per-model methods (get_model, get_model_examples, get_model_pricing) accept either the model's AIR or its catalog slug (the model field returned by list_models).

This is the same data the model picker and pricing pages render from. Listing endpoints accept paginate=True if you want a paginated envelope instead of a flat list.

Utility methods

Beyond inference, the client exposes the platform's utility endpoints. Each is a typed method that takes the same RunOptions second argument as run():

async with Runware() as client:
    # Search the full live model catalog
    models = await client.model_search({"search": "portrait", "architecture": "sdxl", "limit": 10})

    # Store media for reuse as input (URL, data URI, or Base64)
    uploaded = await client.media_storage({"operation": "upload", "media": "https://example.com/photo.jpg"})

    # Account details (credits, limits)
    account = await client.account_management({"operation": "getDetails"})

    # Look up a task you ran earlier, by UUID
    archived = await client.get_task_details({"taskUUID": "abc-123"})

    # Upload a custom model
    await client.model_upload({"category": "checkpoint", "architecture": "sdxl", "format": "safetensors"})

model_search queries the full live catalog, including non-curated models. To browse the curated set as metadata without spending credits, use the content namespace above.

File helpers

file_to_data_uri encodes a local file as a data: URI you can pass anywhere an image input is accepted. It accepts a Path or raw bytes:

from pathlib import Path
from runware import file_to_data_uri

data_uri = file_to_data_uri(Path("photo.jpg"))
await client.media_storage({"operation": "upload", "media": data_uri})

Cancellation and progress

RunOptions.cancel_event accepts an asyncio.Event. Set the event from anywhere and the in-flight call (REST poll, WebSocket subscription, or LLM stream) aborts cleanly and raises a RunwareError with code="aborted".

Cancelling is client-side only. The server keeps processing the task and you are still billed for it. Cancelling just stops the SDK from waiting for the result.

For long-running tasks, two callbacks let you observe a task as it unfolds. on_result fires once per item as it reaches a terminal state, and on_progress fires when an item's progress field (0-100) changes. Only a few long-running models, mostly training, emit progress:

def on_done(item):
    print("Got partial:", item.get("imageUUID") or item.get("videoUUID"))

def on_progress(item):
    print(f"{item.get('progress')}%")

result = await client.run(
    payload,
    RunOptions(on_result=on_done, on_progress=on_progress),
)