Quickstart

Deploy a Python model to a Runware GPU and invoke it, from an empty directory to a running endpoint.

Introduction

One file, one command, one call. This deploys a Python app, waits for it to build, and invokes it.

Before you start

You need a Runware API key, the CLI and the authoring package. The CLI is a separate binary and does not arrive with a pip install.

brew install --cask runware/tap/runware
runware auth login

The package is the other way round, and it is the one your editor needs.

pip install runware-serverless

It declares no dependencies, so it costs you nothing to add. A deploy supplies its own copy either way, which makes this an install for writing the app rather than for running it: it is what resolves the import, checks your handler signatures and lets your own tests call them.

One login binds the CLI to your organization, and everything after that is implicit. Your API key already names your organization, so every command after the login carries it for you.

Write the app

# model.py
from runware_serverless import endpoint, serve


@serve
class Greeter:
    def load(self) -> None:
        self.greeting = "hello"

    @endpoint
    def greet(self, name: str, excited: bool = False) -> dict:
        excited = False if excited is None else excited
        suffix = "!" if excited else "."
        return {"message": f"{self.greeting}, {name}{suffix}"}

load runs once per worker before any request arrives, which is where a real model would be loaded. @endpoint makes greet callable at the path greet. The schema comes from the signature, so name is required and excited is optional.

An omitted optional field arrives as an explicit null, which is why excited applies its own default inside the handler rather than relying on the Python one.

Deploy

runware serverless deploy ./model.py --id my-greeter --gpu-type l40s --wait

A deploy zips and submits the whole directory, so your app can import its own modules and read its own data files.

1
Create

A first deploy with a new --id creates the app. Later deploys with the same id upload a new source and record the next version.

2
Build

The platform builds an image and reads your endpoints out of the source. The app stays initializing until that first build rolls out.

3
Roll

--wait polls until the app is active or failed, so the command finishes when the deploy does rather than when the upload does.

A successful wait is not a running worker. With minWorkers at its default of 0 the app sits scaled to zero until the first invocation arrives, which is also when you first pay for a GPU.

Invoke it

runware serverless apps invoke my-greeter greet --sync -f payload.json

Or over HTTP, which is the same thing without the CLI:

curl -X POST https://api.serverless.runware.ai/v1/apps/my-greeter/invoke-sync/greet \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "taskId": "8f14e45f-ceea-467a-9a2f-6d2b1f2c0b3e",
    "payload": { "name": "world", "excited": true }
  }'

Your fields go inside payload. The taskId is yours to generate, and sending the same one twice returns the task it already names rather than running the work again.

{
  "id": "8f14e45f-ceea-467a-9a2f-6d2b1f2c0b3e",
  "appId": "my-greeter",
  "status": "completed",
  "endpointPath": "greet",
  "output": { "message": "hello, world!" },
  "error": null
}

The first call pays the cold start. Subsequent calls reach a warm worker until the idle window closes.

What to change next

Worker settings are create-time flags. Passing --gpu-type or --max-workers to a deploy on an app that already exists is an error, because those live on the app itself.

runware serverless apps scale my-greeter --max-workers 4

Two more things worth setting early.

Weights belong on a volume. The app runs in a sandbox whose filesystem is part of the checkpointed state, so anything downloaded at runtime and left unmounted is copied into every checkpoint and fetched again on every cold start.

runware serverless deploy ./model.py --id my-greeter --gpu-type l40s \
  --volume /root/.cache/huggingface

Secrets belong in a file, not a flag. A value passed with --env is visible in the process list and recorded in your shell history. Use --env-file for anything that matters.

An app's environment belongs to a version, so setting a variable afterwards records a new version and rolls the workload to pick it up. That is a redeploy of the same image, and a write that does not change the stored value records nothing at all.

Where to go next