Migrating from Runpod

How a Runpod handler maps onto Runware Serverless: the handler becomes a class, the job dict becomes typed parameters, and what stays the same.

Introduction

The part that surprises people coming from Modal or Fal is familiar ground for you. A Runpod endpoint keeps GPU type, worker counts and timeouts in the endpoint settings rather than in your handler file, and so do we. If you started on Flash instead, where the decorator carries the hardware, that part moves out to deploy flags here.

What changes is the handler itself: a single function that unpacks a dictionary becomes a class whose method signatures are the contract.

Side by side

# Runpod
import runpod

pipe = load_my_pipeline()

def handler(job):
    job_input = job["input"]
    prompt = job_input.get("prompt")
    steps = job_input.get("steps", 4)
    return {"image": pipe(prompt, steps).encode()}

runpod.serverless.start({"handler": handler})
# Runware
from runware_serverless import endpoint, serve

@serve
class ImageTools:
    def load(self) -> None:
        self.pipe = load_my_pipeline()

    @endpoint
    def generate(self, prompt: str, steps: int = 4) -> dict:
        steps = 4 if steps is None else steps
        return {"image": self.pipe(prompt, steps).encode()}
runware serverless deploy model.py --id image-tools --gpu-type h100

The start() call goes away. The decorator registers the class, and the platform finds it when the build imports your file.

The job dictionary becomes your signature

This is the real change. Runpod hands you job["input"] and you unpack it yourself, so the shape you accept lives inside the function body. Here the signature is the schema.

RunpodRunware
job["input"]["prompt"]A prompt parameter on the method
.get("steps", 4)A default in the signature, steps: int = 4
A schema is optional, written in your own code and checked inside the handlerThe build publishes it, and the platform rejects a body that does not match

A parameter without a default is required. Runpod's own rp_validator runs inside your handler, after a worker has already woken. Here the schema is closed and published with the build, so an undeclared field is a 422 before any worker is woken.

An omitted optional field arrives as an explicit null, not as a missing key, so a Python default alone does not fire. Apply it inside the handler, as steps = 4 if steps is None else steps above. This is the closest thing to a gotcha in the port.

If you prefer a single object, annotate a parameter with a Pydantic model and it arrives validated as an instance.

One handler becomes many endpoints

A Runpod queue endpoint runs one handler, and its load balancing endpoints get many paths only by you running your own HTTP server. Here one app can expose up to 20 endpoints, each with its own request shape, and the path comes from the method name with underscores turned into hyphens.

@endpoint
def generate(self, prompt: str) -> dict: ...    # answers at generate

@endpoint
def run_upscale(self, image: str) -> dict: ...  # answers at run-upscale

They share one queue and one worker pool, so hardware is still set once on the app. Two endpoints that need different GPUs are two apps.

Loading the model

Runpod leaves loading to you, so most handlers initialize at module scope or lazily on first call. Here it has a name: load runs once before the worker takes traffic, and everything you assign to self is available to every handler.

load is required, even if there is nothing to load, because the build identifies your class by it. And it can run more than once on the same worker: an out-of-memory error at startup retries it, re-instantiating your class, so both __init__ and load have to be safe to run twice.

Configuration

The names change, the model stays the same.

RunpodRunware
Active workers--min-workers
Max workers--max-workers
GPUs per worker--gpus-per-worker, which takes 1, 2, 4 or 8
Idle timeout--idle-ttl
GPU type--gpu-type on the first deploy, plus one --fallback-gpu-type on apps scale
Execution timeoutrequestTimeoutSecs on a container app. A code app keeps the platform ceiling and cannot set one
Job TTLNo equivalent

Runpod lets you rank up to three GPU types so it can fall back during high demand. There is no equivalent here. You name one type, and a worker that cannot get it waits for it. fallbackGpuType is recorded on the app but placement does not substitute it, so plan on the type you pick being the type you get.

All of it is changeable on a live app with apps scale, and a configuration change costs no rebuild. The exception is raising --gpus-per-worker above 1, which needs a redeploy rather than only a new number.

Task ids belong to you

Runpod generates the job id and returns it. Here you generate it and send it, as part of the request envelope.

{
  "taskId": "6b0c2f4e-8d3a-4c1b-9e7f-2a5d8c1b3e60",
  "payload": { "prompt": "a red bicycle" }
}

The id is what makes a retry safe. Sending the same id again returns the task it already names rather than starting a second run, so a lost response can be retried without paying for the work twice. See Invoking an app.

Containers

If you deploy from GitHub today the shape is familiar, because Runpod already builds your image from the Dockerfile in the repo. If you push to a registry instead, that step goes away: Runware always builds and never pulls an image you pushed. Either way you add one file Runpod has no counterpart for, a container.yaml beside the Dockerfile declaring the port and your endpoints.

runware serverless deploy --id my-app --gpu-type h100 --container ./wrapper

See Bringing a container.

What you will miss

Local testing on the code path. Runpod runs your worker locally, either one-shot with python handler.py --test_input '{"input": {...}}' or as a server with python handler.py --rp_serve_api, which answers /runsync on localhost:8000. There is no equivalent for a code app. @serve hands your class back unchanged, so your own tests can import it and call the methods, but nothing exercises the build or the dispatch before you deploy. A container app does not have this gap, because it is your image and you can run the whole thing under Docker first.

Concurrency inside a worker. Runpod's concurrency_modifier lets one worker hold several jobs at once. A Runware worker serves one task at a time, so throughput is the worker count and nothing else. An endpoint that leaned on concurrency rather than on worker count needs a higher --max-workers here to hold the same rate.

Load balancing endpoints. Every invocation here goes through the queue. There is no mode that routes traffic straight to a worker and bypasses it.

Streaming handlers. A handler returns once. There is no incremental result stream.

A GPU priority list. You name one type and a worker waits for it.

Porting checklist

  1. Wrap your handler in a class and put @serve on it
  2. Add a load method, even an empty one, and move both your module-scope initialization and your heavy imports into it, since the build imports your file without those dependencies installed
  3. Turn job["input"] lookups into method parameters, with defaults where you used .get(key, default)
  4. Apply those defaults inside the handler too, since an omitted field arrives as null
  5. Split distinct jobs into separate @endpoint methods instead of branching inside one handler
  6. Drop the runpod.serverless.start(...) call
  7. Move the endpoint settings to flags on your first deploy, and pick the single GPU type you want to run on. Later changes go through apps scale, not through another deploy
  8. Set --max-workers explicitly. It defaults to 1 here, where Runpod defaults to 3
  9. Decide your mount paths and pass them as --volume on that first deploy, because the volume set is fixed when the app is created
  10. Generate a taskId per request and send the envelope