Migrating from Fal
How a fal.App maps onto Runware Serverless: setup and endpoints, machine types, and where scaling configuration moves to.
Introduction
A fal.App ports across almost line for line. You already write a class, you already load weights in a setup method, and you already decorate the methods you want callable. The shape is the same.
What changes is narrow: the hardware leaves the file, the endpoint decorator loses its argument, and fal run goes away.
Side by side
# fal
import fal
from pydantic import BaseModel
class Input(BaseModel):
prompt: str
class MyModel(fal.App):
machine_type = "GPU-H100"
requirements = ["torch", "diffusers"]
def setup(self):
self.pipe = load_my_pipeline()
@fal.endpoint("/generate")
def generate(self, input: Input) -> dict:
return {"image": self.pipe(input.prompt).encode()}# Runware
from pydantic import BaseModel
from runware_serverless import endpoint, serve
class Input(BaseModel):
prompt: str
@serve
class MyModel:
def load(self) -> None:
self.pipe = load_my_pipeline()
@endpoint
def generate(self, input: Input) -> dict:
return {"image": self.pipe(input.prompt).encode()}runware serverless deploy model.py --id my-model --gpu-type h100 \
--requirement torch --requirement diffusers@serve marks a plain class, and the class is the app.
One thing does not port. On Fal the annotated model is the body, so your callers sent {"prompt": "a red bicycle"}. Here each parameter is a member of payload, so the identical signature now asks for {"input": {"prompt": "a red bicycle"}}. Take the model's fields as parameters instead, or keep the model and name the parameter for what it holds.
What leaves the file
machine_type and requirements are class attributes in Fal. Hardware is not expressible in source here at all: --gpu-type is an argument to the deploy and nothing in your file can name it. Dependencies are, just not on the class. Pass them as --requirement, or ship a pyproject.toml and a uv.lock in your source directory and the build reads them.
Fal lets you set these in three places: the class, fal apps scale and the dashboard. Which one wins depends on the parameter. keep_alive, min_concurrency and max_concurrency keep whatever you set outside the code, while machine_type, max_multiplexing and startup_timeout snap back to the class on your next fal deploy.
Here there is one place. Hardware and scaling live on the app, set through the API, the CLI or the dashboard, and changing them never touches your source or your build.
The upside is that a configuration change skips the rebuild, and your scaling survives a code deploy.
The mapping
| Fal | Runware |
class X(fal.App) | @serve on a plain class |
def setup(self) | def load(self), required and named |
@fal.endpoint("/path") | @endpoint, path derived from the method name |
machine_type = "GPU-H100" | --gpu-type h100 |
requirements = [...] | --requirement on the deploy, repeatable |
keep_alive | --idle-ttl |
min_concurrency | --min-workers |
max_concurrency | --max-workers |
ContainerImage | --base-image on a code deploy, handlers unchanged |
Direct Server Mode, exposed_port | A container source: your Dockerfile plus a container.yaml |
fal deploy | runware serverless deploy |
| Revisions | Versions, with apps versions activate to roll back |
fal run | No equivalent |
Fal's concurrency bounds are worker counts here. min_concurrency and max_concurrency bound the number of runners, and --min-workers and --max-workers do the same job.
Fal's max_multiplexing, which lets one runner take several requests at once, has no equivalent, because a Runware worker serves one task at a time. Throughput is the worker count and nothing else, so an app that ran max_multiplexing = 4 needs roughly four times the worker ceiling to hold the same rate.
A list of machine types does not carry across either. Fal treats machine_type = ["GPU-H100", "GPU-A100"] as a pool it tries in order. --fallback-gpu-type records a second type but placement never substitutes it: a worker that cannot get --gpu-type waits rather than starting on the fallback.
Endpoint paths
Fal names the route in the decorator. Runware derives it from the method name, replacing underscores with hyphens, so run_upscale answers at run-upscale.
@endpoint
def generate(self, input: Input) -> dict: ... # answers at generate@fal.endpoint("/") has no direct equivalent, because every path here is a named segment. The closest shape is a class that decorates nothing, which serves its predict at the segment predict.
Because the path comes from the method name, renaming a method renames your public endpoint. There is no argument to pin it. Pick names you are willing to publish.
Pydantic
Your input models transfer unchanged, and they do the same job. A parameter annotated with a model arrives validated, as an instance, and its schema is published as that endpoint's contract. A mismatched body is rejected before a worker is woken.
Output models are not enforced. A return annotation is published as the endpoint's output schema, and nothing checks what your handler actually returned against it.
On Fal the same annotation was enforced, because the endpoint was a FastAPI route and the return type became its response model: fields outside it were dropped and a wrong shape failed the request. Neither happens here, so a response Fal was quietly trimming now goes out whole. Check what your handlers actually return before you point traffic at this.
Loading weights
setup() becomes load(), and the build identifies your class by looking for one that defines it.
load can run more than once on the same worker. It is retried when the GPU is out of memory at startup, and the retry re-instantiates your class, so both __init__ and load have to be safe to run twice.
Containers
Fal gives you two container paths and they land in different places here.
Direct Server Mode becomes a container source: a directory holding your Dockerfile and a container.yaml, which declares the port and the endpoints your server exposes.
ContainerImage does not. Your class keeps its handlers on that path, so it stays a code deploy: pass the image your Dockerfile built on as --base-image and your pip packages as --requirement, and load and your @endpoint methods carry across unchanged. A Dockerfile that installs system packages has to be published as an image you can name there, or move to the container path.
runware serverless deploy --id my-app --gpu-type h100 --container ./wrapperOn that path there is no registry pull. Runware builds the image from your Dockerfile, and the endpoint list comes from container.yaml rather than from decorators. See Bringing a container.
What you will miss
There is no fal run. No throwaway URL on real hardware before you deploy, which means import errors and model-loading failures surface as a failed build rather than in a test run. @serve hands your class back unchanged, so your own tests can import it and call the methods directly, but that does not exercise the build.
There is no Playground. No generated browser UI for an app.
There is no marketplace. Your app is yours to call, and there is nothing to publish it into.
Streaming and realtime endpoints. Fal has SSE streaming and is_websocket=True. Here there are two routes, both POST, plus polling. An app built on either has nothing to migrate to.
teardown() and handle_exit(). There is no shutdown hook, so anything that flushed or closed a connection on the way out has to move into the handler or go away.
Public and shared auth modes. Fal's app_auth makes an app callable without your key, or billed to the caller. Every call here uses your key.
Porting checklist
- Drop
fal.Appas a base class and put@serveon the class - Move
machine_typeto--gpu-typeandrequirementsto--requirement - Rename
setuptoload, and make it safe to run twice - Remove the path argument from each endpoint decorator, then check the method names, since they are now your public paths
- Map
keep_aliveto--idle-ttl, and the concurrency bounds to--min-workersand--max-workers. Set--max-workersexplicitly, because it defaults to1wheremax_concurrencywas unset - Keep your input models. Check what each handler returns, since Fal was validating and trimming it against the output model and nothing here does
- Name your mount paths on
--volumeat create time. Fal gave every runner/datawithHF_HOMEalready pointing into it. Here the paths are yours to declare, and only when the app is created, so getting them wrong means a new app - Update your callers to send the envelope described in Invoking an app