Serverless commands
Deploy, scale, inspect and invoke an app from the terminal, and read what it has cost.
Introduction
Serverless lives inside the existing runware binary as a command group. There is no second tool to install and nothing arrives through pip.
runware serverless --helpThe commands
runware serverless
deploy [file] create an app, or push a new source to one
gpus list GPU types and per-second pricing
open <appId> open the app in the dashboard
usage GPU time and spend, organization-wide
secrets
list · set · remove organization secrets
attach · detach bind a secret to an app
attachments <appId> what is bound to this app
apps
list · show the estate, and one app
endpoints <appId> callable paths, and show one
invoke <appId> <path> send a request
tasks <appId> task history, and show one
workers <appId> live workers, and show one
events <appId> deploy, scaling, audit and error events
versions list · show · activate · delete
builds list · show · delete
env list · set · unset
scale <appId> patch live worker configuration
stop · resume · delete lifecycle
logs <appId> recent logs, or follow them
usage <appId> the same report for one appPaginated listings take --limit and --cursor. Output format is set anywhere with --format or -F, which takes table, json or yaml.
An app printed as json or yaml shows [redacted] for every environment variable value. The names survive, the values are masked, on apps show, deploy, apps scale, stop, resume, delete and version activate. Read the values with apps env list instead.
Deploying
deploy is both create and update, and which one it does depends on whether --id already exists.
runware serverless deploy ./app.py --id my-app --gpu-type h100A first deploy with a new --id creates the app. A later deploy with the same --id uploads a new source, records version N+1, and rolls it out when the build is ready.
Some flags only work on create. --gpu-type, --name, --volume, --env, --env-file and the worker settings apply when the app is created, and passing them to a later deploy is an error rather than a silent no-op. Change workers with apps scale and environment with apps env.
Add --wait to poll until the app is active or failed. A successful wait is not a running worker: with --min-workers at its default of 0 the app sits scaled to zero until the first invocation.
What gets uploaded
You name an entry file, and the whole source directory is packaged, so your entry file can import its own modules and read its own data files. That directory is the working directory unless --src-dir says otherwise, and the entry file has to live inside it.
runware serverless deploy src/app.py --src-dir ~/projects/my-app --id my-app --gpu-type h100Trim the upload with a .runwareignore at the root of the source directory, which takes gitignore syntax.
.gitignore is not consulted. What a project keeps out of version control is a different question from what it ships to a GPU. Either way the upload always skips .env files, .git, __pycache__, .venv, node_modules and the usual build and tool caches.
Container sources
Point --container at a directory whose root holds a Dockerfile and a container.yaml, plus whatever the Dockerfile copies.
runware serverless deploy --id my-app --gpu-type h100 --container ./wrapperRunware builds a hosted image from that archive, so the version records a build rather than an image reference of yours. An invalid container.yaml is rejected at create: 400 if it cannot be parsed, 422 if it parses and breaks a rule.
--container cannot be combined with an entry file, --src-dir, --base-image or --requirement.
Deploy flags
| Flag | Default | What it does |
--id | required | Application ID, immutable after create |
--gpu-type | required on create | See serverless gpus |
--name | --id | Display name, mutable |
--src-dir | working directory | What to package. Code deploys only |
--container | none | Directory holding Dockerfile and container.yaml |
--base-image | python:3.12-slim | Builder base image, Python 3.12 or newer. Code deploys only |
--requirement | none | Extra pip package, repeatable. Code deploys only |
--volume | none | Absolute path backed by persistent storage, repeatable |
--env, --env-file | none | Environment variables, both repeatable |
--min-workers | 0 | Scale floor. 0 allows scale to zero |
--max-workers | 1 | Scale ceiling |
--idle-ttl | 60 | Seconds an idle worker survives |
--scaling-delay | 10 | Cooldown between scaling decisions |
--wait | off | Poll until active or failed |
Environment variables
Supply them on the deploy that creates the app. An app's environment is frozen into the version the deploy creates, and that version is what a worker is rendered from. apps env set on an existing app stores the value and it never reaches a running pod.
Prefer --env-file for anything secret. A value passed as --env is visible in the process list and lands in your shell history.
printf 'HF_TOKEN=%s' "$token" > .env.deploy
runware serverless deploy model.py --id my-app --gpu-type l40s --env-file .env.deployFor credentials that outlive one app, use secrets instead: secrets set once, then secrets attach per app. Attaching rolls the live deployment, which is the difference that matters here.
Scaling
apps scale patches live worker configuration. Omitted flags are left unchanged, so you can move one value without restating the rest.
runware serverless apps scale my-app --max-workers 4 --idle-ttl 120It takes --min-workers, --max-workers, --idle-ttl, --scaling-delay, --min-available-workers, --available-workers-pct, --gpu-type, --fallback-gpu-type and --gpus-per-worker. Changes take effect on the next scaler cycle and the command does not wait for a rollout.
Invoking
runware serverless apps endpoints my-app
runware serverless apps invoke my-app generate -f payload.jsonThe endpoint is a bare path as apps endpoints returns it, and a leading slash is rejected. The default is asynchronous and prints the accepted task id.
| Flag | What it does |
--sync | Use the synchronous route. If the platform wait window expires it polls the task rather than treating expiry as a failure, and never resubmits |
--wait | Poll an async invocation until the task is completed or failed |
-f | Read the JSON payload from a file. Otherwise it reads standard input |
--task-id | Supply the task id instead of generating one |
A task id is sent with every invocation. Resubmitting the same id returns the task it already names rather than starting a second run, so a lost response can be retried without paying twice.
runware serverless apps invoke my-app generate \
--task-id 7c9e6679-7425-40de-944b-e07fc1f90ae7 -f payload.jsonInspecting
runware serverless apps list --status active
runware serverless apps show my-app
runware serverless apps workers my-app
runware serverless apps events my-app --type scaling
runware serverless apps versions list my-app
runware serverless apps builds list my-appapps list also filters with --gpu-type and --query, and orders with --sort. Events filter by --type: deploy, scaling, audit or error. See Monitoring for what to read when.
apps workers shows the active version's workers and takes --version to widen that: an id for one version, or all for every one. An app with nothing pinned lists nothing by default. The cursor replays the version scope along with --state and --status, so changing any of them means starting the pages again.
Logs
runware serverless apps logs my-app --window 6h
runware serverless apps logs my-app --followapps logs prints recent entries over --window: 1h, 6h, 24h, 7d or 30d, default 1h. --limit (1 to 100, default 20) and --cursor page through them.
--sort takes oldest or newest. By default it fetches the newest page and prints it oldest first, so the last line is the most recent. It prints a cursor for the page ahead and the page behind, so you can walk a window from either end. A cursor only works under the sort it was issued with.
--follow prints that page, then streams new entries until you stop it. Entries written while the stream opens or reconnects can be missed or repeated, so a followed stream is not a complete record. --cursor cannot be combined with --follow.
Versions
runware serverless apps versions activate my-app 7Rolling back and rolling forward are the same operation. Nothing is rebuilt, because the version already pinned an image.
Lifecycle
runware serverless apps stop my-app
runware serverless apps resume my-app
runware serverless apps delete my-appA stopped app releases its workers and answers 409 until it is resumed. A source deploy against a stopped app is also a 409, so resume before pushing.
Renaming
runware serverless apps rename my-app "Image generator"Only the display name changes. The app id is immutable, so every URL and every deploy that names it keeps working. The change records a version and leaves your workers alone, so nothing rolls and nothing restarts.
Usage
runware serverless usage --for this-month
runware serverless apps usage my-app --for last-month --group-by dayserverless usage reports GPU time and spend for your whole organization. apps usage <appId> is the same report narrowed to one app. With no flags, both cover the last 24 hours.
| Flag | What it does |
--from | Start of the window, inclusive. RFC 3339, or YYYY-MM-DD read as midnight UTC. Defaults to 24 hours before --to |
--to | End of the window, exclusive. Same formats. Defaults to now |
--for | A UTC calendar range: today, yesterday, this-month or last-month. Cannot be combined with --from or --to |
--gpu-type | Report one GPU type only |
--group-by | Split the figures by app, gpuType, day or coverage, comma-separated |
A window spans at most 31 days. A worker running across either edge counts only for the part inside it, and grouping by day splits at midnight UTC.
The table has one column per grouping, then GPU time, PAYG spend and PAYG-equivalent value. Grouped output ends with a Total row. GPU time is summed per GPU, so two workers running for an hour are two hours. PAYG spend is provisional, not a settled charge. In table output, the window the figures cover is printed to stderr, which keeps it out of anything you pipe. --format json and --format yaml print the API response.
coverage separates time your reserved capacity paid for (commitment) from time billed as you go (payg). PAYG-equivalent value prices all of it at the pay-as-you-go rate, so the gap between that and PAYG spend is what reserved capacity saved. On apps usage, that split reflects how coverage was allocated across the whole organization, because coverage depends on every app's concurrent GPUs.
If any usage in the window cannot be priced, the command fails instead of printing a partial total.