Migrating from Modal

How a Modal app maps onto Runware Serverless: the decorators, the image definition, and where scaling configuration moves to.

Introduction

Most of a Modal app translates line for line. One thing moves, and it is the thing Modal advertises most: Modal puts infrastructure in your source file, and Runware deliberately keeps it out.

Read that difference first, because every other surprise in this page follows from it.

Where the configuration went

In Modal, the GPU and the image are arguments to a decorator. On Runware they are arguments to a deploy.

# Modal
@app.cls(gpu="h100", image=image, min_containers=1)
class Model:
    ...
# Runware
runware serverless deploy model.py --id my-app --gpu-type h100 --min-workers 1

Nothing about hardware or scaling is expressible in the file. The decorator takes no arguments for them, the API rejects an attempt to declare them per endpoint, and a code app has no configuration file.

The reason is that the API owns those values and they are mutable on a live app. You can scale an app from the CLI or the dashboard while it is serving. A source file that also declared the GPU would be stale the moment anyone did, and you would have two places claiming to be true.

A configuration change costs no rebuild, and your scaling survives a code deploy.

Imports move inside

Modal lets you import at module level, or defer with with image.imports():. Here the deferral is not optional: the build imports your file to find your endpoints, and it does that without your dependencies installed, so a module-level import torch fails the build. Move every heavy import into load or into the handler that uses it.

A minimal app, both ways

# Modal
import modal

app = modal.App("image-tools")
image = modal.Image.debian_slim().pip_install("torch", "diffusers", "fastapi[standard]")

@app.cls(gpu="h100", image=image)
class ImageTools:
    @modal.enter()
    def load(self):
        self.pipe = load_my_pipeline()

    @modal.fastapi_endpoint()
    def generate(self, prompt: str):
        return {"image": self.pipe(prompt).encode()}
# Runware
from runware_serverless import endpoint, serve

@serve
class ImageTools:
    def load(self) -> None:
        self.pipe = load_my_pipeline()

    @endpoint
    def generate(self, prompt: str) -> dict:
        return {"image": self.pipe(prompt).encode()}
runware serverless deploy model.py --id image-tools --gpu-type h100 \
  --requirement torch --requirement diffusers

The class is the app: no App object to name, no Image object to build.

--max-workers defaults to 1. Modal's max_containers defaults to no ceiling at all, so an app that scaled freely there lands pinned to a single worker here, with no error anywhere. Set it on the first deploy. The other defaults are --min-workers 0, --idle-ttl 60 and --scaling-delay 10.

The mapping

ModalRunware
modal.App("name")--id on the deploy. The app is created by deploying it
@app.cls(...)@serve, which takes no hardware or scaling arguments, only encode_response
@modal.enter()A load method, required and named
@modal.exit()No equivalent
@modal.fastapi_endpoint()@endpoint, which takes no path
modal.Image...pip_install(...)--base-image and --requirement on the deploy
gpu="h100"--gpu-type h100
min_containers, max_containers--min-workers, --max-workers, and apps scale afterwards
modal.Volume--volume /path. The mount path is the whole identity, there is no name, and it is a node-local cache rather than a distributed volume
modal.Secretsecrets set once, then secrets attach per app
modal deployrunware serverless deploy
modal serveNo equivalent

Endpoints and paths

Modal lets you name a web endpoint through the decorator. Runware derives the path from the method name, replacing underscores with hyphens, so run_upscale answers at run-upscale.

That means renaming a method renames your public endpoint. There is no decorator argument to pin it. Choose method names you are willing to publish, and treat renaming one as a breaking change for your callers.

A class that decorates nothing falls back to serving its predict, so a single-endpoint model needs no decorator at all.

The request body

This is the difference most likely to break a first port. Modal's simple endpoints take query parameters and map them onto your arguments. Runware takes a POST with an envelope, and your fields go inside payload.

{
  "taskId": "6b0c2f4e-8d3a-4c1b-9e7f-2a5d8c1b3e60",
  "payload": { "prompt": "a red bicycle" }
}

taskId is yours to generate, and it makes retries safe: sending the same id again returns the task it already names rather than starting a second run. See Invoking an app.

Your handler still receives the payload fields as keyword arguments, so the shape of the signature carries over.

One thing inside it does not. An omitted optional field arrives as an explicit null, not as a missing key, so a Python default never fires. Re-apply each default in the body, as steps = 4 if steps is None else steps. A positional-only parameter, *args or **kwargs fails the build outright.

Loading weights

@modal.enter() becomes a plain load method, and it is required: the build identifies your class by looking for one that defines load alongside its handlers.

load can run more than once on the same worker. It is retried when the GPU is out of memory at startup, and the retry re-instantiates your class, so both __init__ and load have to be safe to run twice. Modal's @enter runs once per container, so this is a real behavioral difference rather than a rename.

There is no @modal.exit() equivalent. If you rely on one for flushing or cleanup, that logic has to move inside the handler or go away.

Typed requests

One thing that is easier here. Annotate a parameter with a Pydantic model and it arrives validated, as an instance, and the model's schema is published as your endpoint's contract.

from pydantic import BaseModel
from runware_serverless import endpoint, serve

class Transcribe(BaseModel):
    audio_url: str
    language: str | None = None

@endpoint
def transcribe(self, req: Transcribe) -> dict:
    return {"text": self.pipe(req.audio_url)}

A mismatched body is rejected before a worker is woken. See Writing a model.

What you will miss

Worth knowing before you start.

There is no equivalent of modal serve. No live-reloading development loop against the platform. You deploy, and a deploy builds. Iterate against your class locally, since @serve hands it back unchanged and your own tests can call the methods directly.

The CLI validates nothing before it uploads. Importing your model file locally does catch the decorator mistakes, because @endpoint and @serve raise at import for an illegal method name, an argument in the parentheses, or a class with nothing to serve. Everything past that, a missing dependency or a signature the platform cannot bind, surfaces as a failed build rather than an error on your machine.

Concurrency inside a container. Modal's @modal.concurrent(max_inputs=...) lets one container take several requests at once. A Runware worker serves one task at a time, so throughput is the worker count and nothing else. An app that leaned on concurrency needs proportionally more workers here.

No @exit, and no per-endpoint hardware. Every endpoint on an app shares one queue and one worker pool. Two endpoints that genuinely need different GPUs are two apps.

One deployment class per file. Helper classes and request models are fine. What the build refuses is a second class that also defines load alongside endpoints or a predict.

A way to hand over an ASGI or WSGI app. @modal.asgi_app, @modal.wsgi_app and @modal.web_server have no counterpart on the code path, which only ever serves methods on a class. If your Modal deployment is a FastAPI or Flask app, the answer is bringing a container: your image runs the server and a container.yaml declares each route. Note the server receives the inner payload alone, not the envelope around it.

Porting checklist

  1. Strip gpu, image and scaling arguments out of the decorator and hold them for the deploy command
  2. Replace @app.cls(...) with @serve, and @modal.enter() with a load method
  3. Make load and __init__ safe to run twice, since an out-of-memory retry re-instantiates the class
  4. Move every heavy import out of the module level and into load or the handler that uses it
  5. Replace each web endpoint decorator with @endpoint, and check the method names, since they are now your public paths
  6. Move pip_install arguments to --requirement, and the base image to --base-image
  7. Decide your mount paths before the first deploy and pass them as --volume, since anything downloaded at runtime is otherwise refetched on every cold start. Unlike a Modal Volume, the set is fixed when the app is created, so adding one later means a new app
  8. Move @modal.exit() logic into the handler, or drop it
  9. Re-apply every parameter default inside the handler body, since an omitted field arrives as null
  10. Update your callers to send the envelope