Core concepts

The seven nouns Runware Serverless is built from: apps, versions, workers, endpoints, tasks, secrets and volumes.

Introduction

Runware Serverless runs your own model code on Runware. You bring either a Python codebase or a container image, and the platform builds it, scales it to demand, and routes requests to it.

Seven nouns cover the whole product. Everything in the API, the CLI and the SDK is one of them, so the vocabulary here is the vocabulary everywhere else.

App

An app is one workload. It holds your code, the configuration it runs under, and the endpoints it answers on.

Every app has two names and they do different jobs. Getting them the wrong way round is the most common early mistake.

The app ID is the identifier, and you choose it at creation. It cannot be changed afterwards, because everything else in the API resolves against it. Between 6 and 30 characters, lowercase letters, digits and hyphens, starting with a letter and ending with a letter or digit:

^[a-z][a-z0-9-]{4,28}[a-z0-9]$

It has to be unique among your organization's live apps, not globally, so nobody outside your organization can take a name from you. Deleting an app releases its ID for reuse.

The app name is a display label. It is what the dashboard renders and what name sorting orders on, it can repeat across your apps, and you can change it whenever you like. Routing and identity both run on the ID.

Everything else in the API is addressed through the app ID:

POST /v1/apps/{appId}/invoke-sync/{endpointPath}

Version

Every change to an app produces a version, and a version is immutable. It records exactly what to deploy: the image digest the build produced, and the configuration it runs under.

A version is created when you change your code, and also when you change only the configuration. A config-only update carries the previous image forward, so raising your worker ceiling costs a version and reuses the build.

Each version carries a structured diff against the one before it, so you can see what actually changed between two deploys.

Which version serves traffic is app state: you activate one, and every subsequent invocation reaches it. Rolling back is activating an earlier version, which re-applies the image it already built.

Worker

A worker is one running copy of a version, holding one or more GPUs.

The platform starts and stops workers for you. You set the bounds it scales between:

  • minWorkers is the floor. Setting it to 0 allows the app to scale to zero, and you pay for no GPU time while it sits there.
  • maxWorkers is the ceiling, and it is required.
  • idleTtlSecs is how long an idle worker survives before it is reclaimed.
  • scalingDelaySecs damps the scaling decision, so a spike has to persist before it starts workers you then pay for.
  • minAvailableWorkers and availableWorkersPct keep idle workers ahead of demand while work is queued.

A worker serves one task at a time, so throughput is the number of workers and nothing else.

A worker runs on one GPU by default, and gpusPerWorker raises that to 2, 4 or 8. The count is a group size the cluster grants as a unit, and Compute and scaling covers which sizes your GPU type allows.

Scaling to zero has a cost: a request that arrives cold waits for a worker to start and your model to load. Holding minWorkers above zero keeps capacity ready and trades GPU time for latency, which is the main tuning decision this product asks of you.

Worker configuration lives on the app, separate from your source, and can be changed on a running app, which is why a scaling change produces a version and reuses the build.

Endpoint

An endpoint is a callable entry point on your app, and one app can expose several.

An endpoint is identified by its path, a bare lowercase segment with no leading slash, up to 64 characters:

^[a-z]([a-z0-9-]{0,62}[a-z0-9])?$

So generate, transcribe and upscale are valid. /generate is rejected at submit, which catches most people once.

Every invocation is a POST, and the endpoint path is the last segment of the URL.

Where endpoints come from depends on how you deploy. Bring code and the build reads your source, recording one entry per method you marked as an endpoint. Bring a container and you declare them in a config file beside your Dockerfile, since all the builder ever sees is the archive you submit.

An endpoint can declare a schema for its input, and a request that breaks it is then rejected before it reaches a worker, which keeps the failure out of your model. An endpoint that leaves its input open accepts whatever you send it.

Task

Every invocation is a task, whether you called the synchronous route or the asynchronous one.

You generate the task ID yourself and send it with the request:

{
  "taskId": "8f14e45f-ceea-467a-9a2f-6d2b1f2c0b3e",
  "payload": { "prompt": "a red fox in snow" }
}

You choose the task's ID: taskId is a required member of the invocation body, and it is what comes back and what you poll with. Alongside it a task carries a status, the app and endpoint it belongs to, and when it was created. On success the result is in output, on failure the reason is in error. Both are optional, so read the status to tell which you have.

Two things follow from the ID being yours. Sending the same ID twice returns the task it already names, so a request whose response you lost can be retried and still bills once. And you know the ID before the call returns, so a dropped connection still leaves you able to collect the result.

The envelope is strict. taskId and payload are the only two members it takes. Your own fields go inside payload, where they stay clear of every platform field.

The synchronous route waits for the result and returns it. Work that outlives the wait window comes back as a pending task, which you then poll exactly as you would an asynchronous one.

GET /v1/apps/{appId}/tasks/{taskId}

Secret

A secret is credential material your app needs at runtime, such as a Hugging Face token.

Secrets belong to your organization and are attached to the apps that need them. Each attachment resolves to an environment variable name the value is injected as, so your code reads it the way it reads any other. The value reaches the worker at runtime and stays out of the image and out of your application definition.

Plain environment variables are a separate mechanism with their own routes. They share one namespace with secret attachments, so a name used by one is rejected for the other.

Volume

A volume is a persistent directory mounted into your application, for model weights and caches that should survive a cold start and be read again on the next one.

A volume is identified by its mount path and nothing else. There is no name and no separate resource to create:

/models

You can attach up to 30, the paths must be absolute and non-overlapping, and the set is frozen into each version. A worker that already has the weights on disk starts faster than one that has to fetch them, which is where most of a cold start goes.

How they fit together

You create an app and give it a source, either code or a container. The build produces an immutable version with a pinned image and its endpoints recorded.

Deploying activates that version. The platform starts workers up to the bounds you set, mounting your secrets and volumes into each one.

A caller sends a task to an endpoint. The platform validates it, queues it, and routes it to a worker. You read the result back by the task ID you supplied.