Deploy your own models on Runware GPUs
Bring a Python codebase or a container, and Runware builds it, scales it to demand, and routes requests to it.
Introduction
Runware Serverless runs your own model code on Runware. You bring a Python codebase or a container image, and the platform builds it, keeps it running only while there is work, and routes requests to it.
The platform provisions the servers, runs the queues and handles the scaling. You describe what your workload needs and the platform meets it.
When to use it
Reach for Serverless when the model you want to run is yours: a fine-tune, a proprietary pipeline, an architecture nobody hosts, or a chain of steps that has to happen on one machine next to the weights.
If the model you need already exists on Runware, call it through the existing API with a model identifier and a single request. See the platform introduction.
Which path is yours
Two ways in, and the one you take depends on what you already have.
Write a class, mark the methods you want callable, and let the platform build the image for you. You never write a Dockerfile.
from runware_serverless import endpoint, serve
@serve
class ImageTools:
def load(self) -> None:
self.pipe = load_my_pipeline()
@endpoint
def generate(self, prompt: str, steps: int = 4) -> dict:
return {"image": self.pipe(prompt, steps).encode()}The build reads your handler signatures and derives the request schema from them, so a malformed request is rejected before it reaches your code. See Writing a model.
Bring your own Dockerfile and the HTTP server inside it. Because the builder cannot see inside your image, you declare the endpoints alongside it in a small configuration file.
configVersion: 1
port: 8080
probes:
readiness: /ready
liveness: /health
endpoints:
- path: generate
input:
type: object
required: [prompt]
properties:
prompt: { type: string }Everything after the build is identical to the code path. See Bringing a container.
How a deployment works
An app is one workload. You choose its ID at creation and it is permanent, because everything else resolves against it.
The platform packages your source into an image and records the endpoints it found. The result is an immutable version with the image digest pinned.
Activating a version points traffic at it. Rolling back is activating an earlier one, with no rebuild.
Send a task to an endpoint and read the result. The platform starts workers when requests arrive and releases them when the traffic stops.
What you configure
You set the shape of the compute and the platform stays inside it. A GPU type, how many GPUs each worker gets, and minWorkers and maxWorkers for how many run at once, plus how long an idle worker survives and how many requests one worker takes at a time.
Setting minWorkers to zero means you pay for no GPU while nothing is arriving, at the cost of a cold start on the next request. Holding it above zero buys latency with GPU time. That trade is yours to make and it is the main one this product asks of you.
Scaling configuration lives outside your source. One command changes it on a running app with no rebuild, so you can tune it during an incident.
Getting started
- Sign up for a Runware account and generate an API key.
- Install the CLI and authenticate with
runware auth login. - Follow the quickstart to deploy and invoke your first app.
- Read Core concepts for the seven nouns everything else is built from.