Python SDK

Define an application in Python and deploy it.

Introduction

A code app is one class in one file, and the whole authoring surface is two decorators and one method.

pip install runware-serverless

The package declares no dependencies and exports exactly serve and endpoint. A deploy supplies its own copy, so installing it is what gives you a resolved import and a checked signature while you write, rather than something a deployment depends on.

from runware_serverless import endpoint, serve

@serve
class ImageTools:
    def load(self) -> None:
        self.pipeline = load_pipeline()

    @endpoint
    def generate(self, prompt: str, steps: int = 4) -> dict[str, object]:
        return {"image": self.pipeline(prompt, steps)}

Writing a model walks through it. This page is the reference.

You do not install this package to deploy. The platform installs it for the build and again in every worker, so the version your app runs against is always the one the platform serves it with.

The class

One decorated class per model file. The build reads the file, finds the class and refuses a file with more than one.

To be that class, it must be defined in the model file itself, carry a callable load, and declare at least one endpoint: a method marked with @endpoint, or a callable predict when no method is marked. A class with neither fails the build.

serve

@serve registers the class as the app the worker serves, and returns the class unchanged, so it stays importable, testable and subclassable. It resolves the endpoints as the file is imported, so a class with nothing to serve fails the build, before any request reaches it.

Used bare, or called with either option:

@serve(encode_response=to_base64_png)
class Flux:
    ...
OptionWhat it does
encode_responseTakes the list of handler results and returns what goes on the wire. The seam for a response your signature does not describe, so the published contract stays truthful
free_bytesCalled to reclaim VRAM from this worker, in place of the default. A model that can drop something on demand says so here

endpoint

@endpoint marks a method as callable. The path is the method name with underscores turned into hyphens, so run_upscale serves run-upscale. The decorator takes no arguments, and a method name that yields an illegal path fails the build.

A path is lowercase, starts with a letter, and holds letters, digits and hyphens, up to 64 characters. An app serves at most 20 endpoints, and endpoints are inherited: a class built on a shared base serves the base's endpoints too, and overriding one keeps its path.

The decorated method is returned unwrapped, so its signature, annotations and defaults stay exactly as written, for the platform and for your editor.

load

load runs once the worker's container starts and before it reports itself ready, which is where weights belong. It can run more than once: a load that runs out of VRAM is retried with a larger request, and the retry re-instantiates your class, so both __init__ and load have to be safe to repeat.

What the platform reads from your code

Your handler's parameters become the endpoint's request schema, and the platform validates every request against it before your code runs. A parameter annotated with a Pydantic model contributes that model's own schema, and arrives as an instance.

A return annotation is published as the endpoint's response schema and is not enforced, so what your handler returns goes out as it is. See Writing a model.

What lives outside your code

GPU type, worker counts and idle TTL belong to the app rather than to the file: you set them when you create or update the app, from the CLI, the API or the dashboard. See Compute and scaling.