Migrating from Runpod
How a Runpod handler maps onto Runware Serverless: the handler becomes a class, the job dict becomes typed parameters, and what stays the same.
Introduction
The part that surprises people coming from Modal or Fal is familiar ground for you. A Runpod endpoint keeps GPU type, worker counts and timeouts in the endpoint settings rather than in your handler file, and so do we. If you started on Flash instead, where the decorator carries the hardware, that part moves out to deploy flags here.
What changes is the handler itself: a single function that unpacks a dictionary becomes a class whose method signatures are the contract.
Side by side
# Runpod
import runpod
pipe = load_my_pipeline()
def handler(job):
job_input = job["input"]
prompt = job_input.get("prompt")
steps = job_input.get("steps", 4)
return {"image": pipe(prompt, steps).encode()}
runpod.serverless.start({"handler": handler})# Runware
from runware_serverless import endpoint, serve
@serve
class ImageTools:
def load(self) -> None:
self.pipe = load_my_pipeline()
@endpoint
def generate(self, prompt: str, steps: int = 4) -> dict:
steps = 4 if steps is None else steps
return {"image": self.pipe(prompt, steps).encode()}runware serverless deploy model.py --id image-tools --gpu-type h100The start() call goes away. The decorator registers the class, and the platform finds it when the build imports your file.
The job dictionary becomes your signature
This is the real change. Runpod hands you job["input"] and you unpack it yourself, so the shape you accept lives inside the function body. Here the signature is the schema.
| Runpod | Runware |
job["input"]["prompt"] | A prompt parameter on the method |
.get("steps", 4) | A default in the signature, steps: int = 4 |
| A schema is optional, written in your own code and checked inside the handler | The build publishes it, and the platform rejects a body that does not match |
A parameter without a default is required. Runpod's own rp_validator runs inside your handler, after a worker has already woken. Here the schema is closed and published with the build, so an undeclared field is a 422 before any worker is woken.
An omitted optional field arrives as an explicit null, not as a missing key, so a Python default alone does not fire. Apply it inside the handler, as steps = 4 if steps is None else steps above. This is the closest thing to a gotcha in the port.
If you prefer a single object, annotate a parameter with a Pydantic model and it arrives validated as an instance.
One handler becomes many endpoints
A Runpod queue endpoint runs one handler, and its load balancing endpoints get many paths only by you running your own HTTP server. Here one app can expose up to 20 endpoints, each with its own request shape, and the path comes from the method name with underscores turned into hyphens.
@endpoint
def generate(self, prompt: str) -> dict: ... # answers at generate
@endpoint
def run_upscale(self, image: str) -> dict: ... # answers at run-upscaleThey share one queue and one worker pool, so hardware is still set once on the app. Two endpoints that need different GPUs are two apps.
Loading the model
Runpod leaves loading to you, so most handlers initialize at module scope or lazily on first call. Here it has a name: load runs once before the worker takes traffic, and everything you assign to self is available to every handler.
load is required, even if there is nothing to load, because the build identifies your class by it. And it can run more than once on the same worker: an out-of-memory error at startup retries it, re-instantiating your class, so both __init__ and load have to be safe to run twice.
Configuration
The names change, the model stays the same.
| Runpod | Runware |
| Active workers | --min-workers |
| Max workers | --max-workers |
| GPUs per worker | --gpus-per-worker, which takes 1, 2, 4 or 8 |
| Idle timeout | --idle-ttl |
| GPU type | --gpu-type on the first deploy, plus one --fallback-gpu-type on apps scale |
| Execution timeout | requestTimeoutSecs on a container app. A code app keeps the platform ceiling and cannot set one |
| Job TTL | No equivalent |
Runpod lets you rank up to three GPU types so it can fall back during high demand. There is no equivalent here. You name one type, and a worker that cannot get it waits for it. fallbackGpuType is recorded on the app but placement does not substitute it, so plan on the type you pick being the type you get.
All of it is changeable on a live app with apps scale, and a configuration change costs no rebuild. The exception is raising --gpus-per-worker above 1, which needs a redeploy rather than only a new number.
Task ids belong to you
Runpod generates the job id and returns it. Here you generate it and send it, as part of the request envelope.
{
"taskId": "6b0c2f4e-8d3a-4c1b-9e7f-2a5d8c1b3e60",
"payload": { "prompt": "a red bicycle" }
}The id is what makes a retry safe. Sending the same id again returns the task it already names rather than starting a second run, so a lost response can be retried without paying for the work twice. See Invoking an app.
Containers
If you deploy from GitHub today the shape is familiar, because Runpod already builds your image from the Dockerfile in the repo. If you push to a registry instead, that step goes away: Runware always builds and never pulls an image you pushed. Either way you add one file Runpod has no counterpart for, a container.yaml beside the Dockerfile declaring the port and your endpoints.
runware serverless deploy --id my-app --gpu-type h100 --container ./wrapperSee Bringing a container.
What you will miss
Local testing on the code path. Runpod runs your worker locally, either one-shot with python handler.py --test_input '{"input": {...}}' or as a server with python handler.py --rp_serve_api, which answers /runsync on localhost:8000. There is no equivalent for a code app. @serve hands your class back unchanged, so your own tests can import it and call the methods, but nothing exercises the build or the dispatch before you deploy. A container app does not have this gap, because it is your image and you can run the whole thing under Docker first.
Concurrency inside a worker. Runpod's concurrency_modifier lets one worker hold several jobs at once. A Runware worker serves one task at a time, so throughput is the worker count and nothing else. An endpoint that leaned on concurrency rather than on worker count needs a higher --max-workers here to hold the same rate.
Load balancing endpoints. Every invocation here goes through the queue. There is no mode that routes traffic straight to a worker and bypasses it.
Streaming handlers. A handler returns once. There is no incremental result stream.
A GPU priority list. You name one type and a worker waits for it.
Porting checklist
- Wrap your handler in a class and put
@serveon it - Add a
loadmethod, even an empty one, and move both your module-scope initialization and your heavy imports into it, since the build imports your file without those dependencies installed - Turn
job["input"]lookups into method parameters, with defaults where you used.get(key, default) - Apply those defaults inside the handler too, since an omitted field arrives as
null - Split distinct jobs into separate
@endpointmethods instead of branching inside one handler - Drop the
runpod.serverless.start(...)call - Move the endpoint settings to flags on your first deploy, and pick the single GPU type you want to run on. Later changes go through
apps scale, not through another deploy - Set
--max-workersexplicitly. It defaults to1here, where Runpod defaults to3 - Decide your mount paths and pass them as
--volumeon that first deploy, because the volume set is fixed when the app is created - Generate a
taskIdper request and send the envelope