Migrating from Fal

How a fal.App maps onto Runware Serverless: setup and endpoints, machine types, and where scaling configuration moves to.

Introduction

A fal.App ports across almost line for line. You already write a class, you already load weights in a setup method, and you already decorate the methods you want callable. The shape is the same.

What changes is narrow: the hardware leaves the file, the endpoint decorator loses its argument, and fal run goes away.

Side by side

# fal
import fal
from pydantic import BaseModel

class Input(BaseModel):
    prompt: str

class MyModel(fal.App):
    machine_type = "GPU-H100"
    requirements = ["torch", "diffusers"]

    def setup(self):
        self.pipe = load_my_pipeline()

    @fal.endpoint("/generate")
    def generate(self, input: Input) -> dict:
        return {"image": self.pipe(input.prompt).encode()}
# Runware
from pydantic import BaseModel
from runware_serverless import endpoint, serve

class Input(BaseModel):
    prompt: str

@serve
class MyModel:
    def load(self) -> None:
        self.pipe = load_my_pipeline()

    @endpoint
    def generate(self, input: Input) -> dict:
        return {"image": self.pipe(input.prompt).encode()}
runware serverless deploy model.py --id my-model --gpu-type h100 \
  --requirement torch --requirement diffusers

@serve marks a plain class, and the class is the app.

One thing does not port. On Fal the annotated model is the body, so your callers sent {"prompt": "a red bicycle"}. Here each parameter is a member of payload, so the identical signature now asks for {"input": {"prompt": "a red bicycle"}}. Take the model's fields as parameters instead, or keep the model and name the parameter for what it holds.

What leaves the file

machine_type and requirements are class attributes in Fal. Hardware is not expressible in source here at all: --gpu-type is an argument to the deploy and nothing in your file can name it. Dependencies are, just not on the class. Pass them as --requirement, or ship a pyproject.toml and a uv.lock in your source directory and the build reads them.

Fal lets you set these in three places: the class, fal apps scale and the dashboard. Which one wins depends on the parameter. keep_alive, min_concurrency and max_concurrency keep whatever you set outside the code, while machine_type, max_multiplexing and startup_timeout snap back to the class on your next fal deploy.

Here there is one place. Hardware and scaling live on the app, set through the API, the CLI or the dashboard, and changing them never touches your source or your build.

The upside is that a configuration change skips the rebuild, and your scaling survives a code deploy.

The mapping

FalRunware
class X(fal.App)@serve on a plain class
def setup(self)def load(self), required and named
@fal.endpoint("/path")@endpoint, path derived from the method name
machine_type = "GPU-H100"--gpu-type h100
requirements = [...]--requirement on the deploy, repeatable
keep_alive--idle-ttl
min_concurrency--min-workers
max_concurrency--max-workers
ContainerImage--base-image on a code deploy, handlers unchanged
Direct Server Mode, exposed_portA container source: your Dockerfile plus a container.yaml
fal deployrunware serverless deploy
RevisionsVersions, with apps versions activate to roll back
fal runNo equivalent

Fal's concurrency bounds are worker counts here. min_concurrency and max_concurrency bound the number of runners, and --min-workers and --max-workers do the same job.

Fal's max_multiplexing, which lets one runner take several requests at once, has no equivalent, because a Runware worker serves one task at a time. Throughput is the worker count and nothing else, so an app that ran max_multiplexing = 4 needs roughly four times the worker ceiling to hold the same rate.

A list of machine types does not carry across either. Fal treats machine_type = ["GPU-H100", "GPU-A100"] as a pool it tries in order. --fallback-gpu-type records a second type but placement never substitutes it: a worker that cannot get --gpu-type waits rather than starting on the fallback.

Endpoint paths

Fal names the route in the decorator. Runware derives it from the method name, replacing underscores with hyphens, so run_upscale answers at run-upscale.

@endpoint
def generate(self, input: Input) -> dict: ...   # answers at generate

@fal.endpoint("/") has no direct equivalent, because every path here is a named segment. The closest shape is a class that decorates nothing, which serves its predict at the segment predict.

Because the path comes from the method name, renaming a method renames your public endpoint. There is no argument to pin it. Pick names you are willing to publish.

Pydantic

Your input models transfer unchanged, and they do the same job. A parameter annotated with a model arrives validated, as an instance, and its schema is published as that endpoint's contract. A mismatched body is rejected before a worker is woken.

Output models are not enforced. A return annotation is published as the endpoint's output schema, and nothing checks what your handler actually returned against it.

On Fal the same annotation was enforced, because the endpoint was a FastAPI route and the return type became its response model: fields outside it were dropped and a wrong shape failed the request. Neither happens here, so a response Fal was quietly trimming now goes out whole. Check what your handlers actually return before you point traffic at this.

Loading weights

setup() becomes load(), and the build identifies your class by looking for one that defines it.

load can run more than once on the same worker. It is retried when the GPU is out of memory at startup, and the retry re-instantiates your class, so both __init__ and load have to be safe to run twice.

Containers

Fal gives you two container paths and they land in different places here.

Direct Server Mode becomes a container source: a directory holding your Dockerfile and a container.yaml, which declares the port and the endpoints your server exposes.

ContainerImage does not. Your class keeps its handlers on that path, so it stays a code deploy: pass the image your Dockerfile built on as --base-image and your pip packages as --requirement, and load and your @endpoint methods carry across unchanged. A Dockerfile that installs system packages has to be published as an image you can name there, or move to the container path.

runware serverless deploy --id my-app --gpu-type h100 --container ./wrapper

On that path there is no registry pull. Runware builds the image from your Dockerfile, and the endpoint list comes from container.yaml rather than from decorators. See Bringing a container.

What you will miss

There is no fal run. No throwaway URL on real hardware before you deploy, which means import errors and model-loading failures surface as a failed build rather than in a test run. @serve hands your class back unchanged, so your own tests can import it and call the methods directly, but that does not exercise the build.

There is no Playground. No generated browser UI for an app.

There is no marketplace. Your app is yours to call, and there is nothing to publish it into.

Streaming and realtime endpoints. Fal has SSE streaming and is_websocket=True. Here there are two routes, both POST, plus polling. An app built on either has nothing to migrate to.

teardown() and handle_exit(). There is no shutdown hook, so anything that flushed or closed a connection on the way out has to move into the handler or go away.

Public and shared auth modes. Fal's app_auth makes an app callable without your key, or billed to the caller. Every call here uses your key.

Porting checklist

  1. Drop fal.App as a base class and put @serve on the class
  2. Move machine_type to --gpu-type and requirements to --requirement
  3. Rename setup to load, and make it safe to run twice
  4. Remove the path argument from each endpoint decorator, then check the method names, since they are now your public paths
  5. Map keep_alive to --idle-ttl, and the concurrency bounds to --min-workers and --max-workers. Set --max-workers explicitly, because it defaults to 1 where max_concurrency was unset
  6. Keep your input models. Check what each handler returns, since Fal was validating and trimming it against the output model and nothing here does
  7. Name your mount paths on --volume at create time. Fal gave every runner /data with HF_HOME already pointing into it. Here the paths are yours to declare, and only when the app is created, so getting them wrong means a new app
  8. Update your callers to send the envelope described in Invoking an app