Serverless commands

Deploy, scale, inspect and invoke an app from the terminal, and read what it has cost.

Introduction

Serverless lives inside the existing runware binary as a command group. There is no second tool to install and nothing arrives through pip.

runware serverless --help

The commands

runware serverless
  deploy [file]              create an app, or push a new source to one
  gpus                       list GPU types and per-second pricing
  open <appId>               open the app in the dashboard
  usage                      GPU time and spend, organization-wide
  secrets
    list · set · remove      organization secrets
    attach · detach          bind a secret to an app
    attachments <appId>      what is bound to this app
  apps
    list · show              the estate, and one app
    endpoints <appId>        callable paths, and show one
    invoke <appId> <path>    send a request
    tasks <appId>            task history, and show one
    workers <appId>          live workers, and show one
    events <appId>           deploy, scaling, audit and error events
    versions                 list · show · activate · delete
    builds                   list · show · delete
    env                      list · set · unset
    scale <appId>            patch live worker configuration
    stop · resume · delete   lifecycle
    logs <appId>             recent logs, or follow them
    usage <appId>            the same report for one app

Paginated listings take --limit and --cursor. Output format is set anywhere with --format or -F, which takes table, json or yaml.

An app printed as json or yaml shows [redacted] for every environment variable value. The names survive, the values are masked, on apps show, deploy, apps scale, stop, resume, delete and version activate. Read the values with apps env list instead.

Deploying

deploy is both create and update, and which one it does depends on whether --id already exists.

runware serverless deploy ./app.py --id my-app --gpu-type h100

A first deploy with a new --id creates the app. A later deploy with the same --id uploads a new source, records version N+1, and rolls it out when the build is ready.

Some flags only work on create. --gpu-type, --name, --volume, --env, --env-file and the worker settings apply when the app is created, and passing them to a later deploy is an error rather than a silent no-op. Change workers with apps scale and environment with apps env.

Add --wait to poll until the app is active or failed. A successful wait is not a running worker: with --min-workers at its default of 0 the app sits scaled to zero until the first invocation.

What gets uploaded

You name an entry file, and the whole source directory is packaged, so your entry file can import its own modules and read its own data files. That directory is the working directory unless --src-dir says otherwise, and the entry file has to live inside it.

runware serverless deploy src/app.py --src-dir ~/projects/my-app --id my-app --gpu-type h100

Trim the upload with a .runwareignore at the root of the source directory, which takes gitignore syntax.

.gitignore is not consulted. What a project keeps out of version control is a different question from what it ships to a GPU. Either way the upload always skips .env files, .git, __pycache__, .venv, node_modules and the usual build and tool caches.

Container sources

Point --container at a directory whose root holds a Dockerfile and a container.yaml, plus whatever the Dockerfile copies.

runware serverless deploy --id my-app --gpu-type h100 --container ./wrapper

Runware builds a hosted image from that archive, so the version records a build rather than an image reference of yours. An invalid container.yaml is rejected at create: 400 if it cannot be parsed, 422 if it parses and breaks a rule.

--container cannot be combined with an entry file, --src-dir, --base-image or --requirement.

Deploy flags

FlagDefaultWhat it does
--idrequiredApplication ID, immutable after create
--gpu-typerequired on createSee serverless gpus
--name--idDisplay name, mutable
--src-dirworking directoryWhat to package. Code deploys only
--containernoneDirectory holding Dockerfile and container.yaml
--base-imagepython:3.12-slimBuilder base image, Python 3.12 or newer. Code deploys only
--requirementnoneExtra pip package, repeatable. Code deploys only
--volumenoneAbsolute path backed by persistent storage, repeatable
--env, --env-filenoneEnvironment variables, both repeatable
--min-workers0Scale floor. 0 allows scale to zero
--max-workers1Scale ceiling
--idle-ttl60Seconds an idle worker survives
--scaling-delay10Cooldown between scaling decisions
--waitoffPoll until active or failed

Environment variables

Supply them on the deploy that creates the app. An app's environment is frozen into the version the deploy creates, and that version is what a worker is rendered from. apps env set on an existing app stores the value and it never reaches a running pod.

Prefer --env-file for anything secret. A value passed as --env is visible in the process list and lands in your shell history.

printf 'HF_TOKEN=%s' "$token" > .env.deploy
runware serverless deploy model.py --id my-app --gpu-type l40s --env-file .env.deploy

For credentials that outlive one app, use secrets instead: secrets set once, then secrets attach per app. Attaching rolls the live deployment, which is the difference that matters here.

Scaling

apps scale patches live worker configuration. Omitted flags are left unchanged, so you can move one value without restating the rest.

runware serverless apps scale my-app --max-workers 4 --idle-ttl 120

It takes --min-workers, --max-workers, --idle-ttl, --scaling-delay, --min-available-workers, --available-workers-pct, --gpu-type, --fallback-gpu-type and --gpus-per-worker. Changes take effect on the next scaler cycle and the command does not wait for a rollout.

Invoking

runware serverless apps endpoints my-app
runware serverless apps invoke my-app generate -f payload.json

The endpoint is a bare path as apps endpoints returns it, and a leading slash is rejected. The default is asynchronous and prints the accepted task id.

FlagWhat it does
--syncUse the synchronous route. If the platform wait window expires it polls the task rather than treating expiry as a failure, and never resubmits
--waitPoll an async invocation until the task is completed or failed
-fRead the JSON payload from a file. Otherwise it reads standard input
--task-idSupply the task id instead of generating one

A task id is sent with every invocation. Resubmitting the same id returns the task it already names rather than starting a second run, so a lost response can be retried without paying twice.

runware serverless apps invoke my-app generate \
  --task-id 7c9e6679-7425-40de-944b-e07fc1f90ae7 -f payload.json

Inspecting

runware serverless apps list --status active
runware serverless apps show my-app
runware serverless apps workers my-app
runware serverless apps events my-app --type scaling
runware serverless apps versions list my-app
runware serverless apps builds list my-app

apps list also filters with --gpu-type and --query, and orders with --sort. Events filter by --type: deploy, scaling, audit or error. See Monitoring for what to read when.

apps workers shows the active version's workers and takes --version to widen that: an id for one version, or all for every one. An app with nothing pinned lists nothing by default. The cursor replays the version scope along with --state and --status, so changing any of them means starting the pages again.

Logs

runware serverless apps logs my-app --window 6h
runware serverless apps logs my-app --follow

apps logs prints recent entries over --window: 1h, 6h, 24h, 7d or 30d, default 1h. --limit (1 to 100, default 20) and --cursor page through them.

--sort takes oldest or newest. By default it fetches the newest page and prints it oldest first, so the last line is the most recent. It prints a cursor for the page ahead and the page behind, so you can walk a window from either end. A cursor only works under the sort it was issued with.

--follow prints that page, then streams new entries until you stop it. Entries written while the stream opens or reconnects can be missed or repeated, so a followed stream is not a complete record. --cursor cannot be combined with --follow.

Versions

runware serverless apps versions activate my-app 7

Rolling back and rolling forward are the same operation. Nothing is rebuilt, because the version already pinned an image.

Lifecycle

runware serverless apps stop my-app
runware serverless apps resume my-app
runware serverless apps delete my-app

A stopped app releases its workers and answers 409 until it is resumed. A source deploy against a stopped app is also a 409, so resume before pushing.

Renaming

runware serverless apps rename my-app "Image generator"

Only the display name changes. The app id is immutable, so every URL and every deploy that names it keeps working. The change records a version and leaves your workers alone, so nothing rolls and nothing restarts.

Usage

runware serverless usage --for this-month
runware serverless apps usage my-app --for last-month --group-by day

serverless usage reports GPU time and spend for your whole organization. apps usage <appId> is the same report narrowed to one app. With no flags, both cover the last 24 hours.

FlagWhat it does
--fromStart of the window, inclusive. RFC 3339, or YYYY-MM-DD read as midnight UTC. Defaults to 24 hours before --to
--toEnd of the window, exclusive. Same formats. Defaults to now
--forA UTC calendar range: today, yesterday, this-month or last-month. Cannot be combined with --from or --to
--gpu-typeReport one GPU type only
--group-bySplit the figures by app, gpuType, day or coverage, comma-separated

A window spans at most 31 days. A worker running across either edge counts only for the part inside it, and grouping by day splits at midnight UTC.

The table has one column per grouping, then GPU time, PAYG spend and PAYG-equivalent value. Grouped output ends with a Total row. GPU time is summed per GPU, so two workers running for an hour are two hours. PAYG spend is provisional, not a settled charge. In table output, the window the figures cover is printed to stderr, which keeps it out of anything you pipe. --format json and --format yaml print the API response.

coverage separates time your reserved capacity paid for (commitment) from time billed as you go (payg). PAYG-equivalent value prices all of it at the pay-as-you-go rate, so the gap between that and PAYG spend is what reserved capacity saved. On apps usage, that split reflects how coverage was allocated across the whole organization, because coverage depends on every app's concurrent GPUs.

If any usage in the window cannot be priced, the command fails instead of printing a partial total.