Apps

Create an app, change its configuration, deploy a version, and control its lifecycle.

Introduction

An app is the unit you create, configure and invoke. Everything else on this API hangs beneath one.

Creating one names the source it runs, so the archive goes up before the app exists. Source uploads covers that flow and the sourceId it produces.

appId is chosen by you at creation and cannot be changed afterwards. It is released again once the app it named reaches deleted, so a name can be reused by a new app that shares nothing with the old one.

List apps

GETapi.serverless.runware.ai/v1/apps

Returns a page of the organization's apps. Filters combine with AND. Soft-deleted apps are excluded unless status=deleted is requested explicitly. Favorited apps appear before non-favorited apps, with the selected ordering applied within each group. A cursor is only valid for the sort and filters it was issued under. Reusing one across a different ordering or filter set returns 400.

Request

Query

limit

integerint32min: 1max: 100default: 20

Maximum number of items to return.

cursor

string

Opaque pagination cursor returned by a previous call, as nextCursor or, on the operations that offer one, prevCursor.

status

string

Return only apps in this status.

Allowed values7 values

q

stringmin: 1max: 100

Case-insensitive substring match against appName and appId. An app matching either is returned.

gpuType

stringmin: 1max: 64

Return only apps whose worker configuration requests this GPU type. Matched against configuration.gpuType only, not fallbackGpuType. Must be a code from GET /v1/gpu-types. An unknown code returns 422.

sort

stringdefault: createdAt

Ordering for listApps. Favorited apps appear before non-favorited apps, and the selected ordering applies within each group. Every ordering is total (ties broken by appId), so a page is reproducible and its cursor stable. - createdAt: newest first. The default. - name: appName A–Z, case-insensitive.

Allowed values2 values

Response

nextCursor

stringnullable

Cursor for the next page. Null when there are no more items.

data

object[]
Array items13 properties each

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType
stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value
gpuType
stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType
stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker
integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers
integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers
integerrequiredint32min: 1
minAvailableWorkers
integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct
integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs
integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs
integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs
integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs
integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties
state
stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values
reason
stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values
since
string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt
string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties
activeWorkers
integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount
integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt
string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h
integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h
numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h
numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth
integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type
stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value
metadata
objectnullable

Optional opaque metadata associated with the secret.

createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key
stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
400The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
401Missing or invalid credentials
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

Create an app

POSTapi.serverless.runware.ai/v1/apps

Creates an app together with its worker configuration, environment variables and endpoints, and records version 1, the immutable description of what to deploy. The app starts in initializing, and what happens next depends on the app source type: - code source: the codebase is submitted to the build pipeline. Once the image is built and workers become healthy the app transitions to active and activeVersionId points at that version. If the build, validation, or rollout fails the app is marked failed. - container source: the submitted zip (wrapper Dockerfile + container.yaml) goes through the same build pipeline. The wrapper image is built, published and deployed, so the version carries a buildId and the app follows the same lifecycle as a code source. An invalid container.yaml rejects the create before any build capacity is spent: 400 where the document could not be parsed at all, 422 where it parsed and broke a rule. activeVersionId is null until a rollout completes: a version records what should run, and only a finished deploy says what does. secrets attaches organization secrets that already exist. It is the app's initial attachment set, so the first rollout carries their values into the worker. This route does not create a secret. Use POST /v1/secrets first. A name that is unknown to the organization, or that is not active, returns 404. A name that collides with a key in environmentVariables, a repeated name and a set that goes past the binding limit each return 422. The whole set is checked before any build capacity is spent.

Creating an app does not make it callable. The first build has to finish and roll out, and until it does the app stays initializing and answers 409.

Request

Body

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired
Properties12 properties

computeType

stringdefault: gpu

GPU is the only supported compute type. Omitting the field selects GPU. CPU workloads are not supported. A request that names cpu is rejected with 422 before a build or deploy starts.

Allowed values1 value

gpuType

stringrequiredmin: 1max: 64

GPU type the workers run on. Required: omitting it (or sending null) is a 422 before a build or deploy starts, because a GPU app with no type is unpinned and the deployer would render NVIDIA defaults. Must match an id returned by GET /v1/gpu-types that currently has admitted capacity.

fallbackGpuType

stringnullablemax: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it, so it is not failover you can rely on today. Omit, send JSON null, or send an empty string for no fallback. Forms bind an unselected dropdown as "", which is not a GpuTypeId. A non-empty value must be an active catalog code. Unlike gpuType this is an existence check only: it does not require admitted capacity, so the code may not appear in the customer GET /v1/gpu-types list.

gpusPerWorker

integerint32default: 1

GPUs granted to one worker pod. A worker holds its GPUs as one group the cluster grants indivisibly, so the count is one of the advertised group sizes rather than any number in a range. The value must also be a group size admitted by the cluster backing the chosen gpuType: a count above 1 that cluster does not grant is rejected with a 422 naming /configuration/gpusPerWorker, since the pod could never be scheduled. A value above 1 requires an image built after the multi-GPU worker entrypoint. Older images serve a single rank while holding every granted GPU.

Allowed values4 values

minWorkers

integerint32min: 0default: 0

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand. Omit, or send null or 0, for no buffer. Must be lower than maxWorkers. The buffer applies only while the queue is non-empty, so it does not stop an idle app scaling to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Omit, or send null or 0, for no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty, so it does not stop an idle app scaling to minWorkers.

idleTtlSecs

integerrequiredint32min: 1

scalingDelaySecs

integerrequiredint32min: 1

requestTimeoutSecs

integerint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only. Omit on a container app to take 60s. A code app may still send or omit this field: omit stores 600s for the console, and neither value is applied to MLflow reads, which keep the platform ceiling. A later container.yaml deploy overwrites the live value with that document's timeouts.requestSeconds. Bounds the worker, not the synchronous HTTP wait.

startupTimeoutSecs

integerint32min: 1max: 2400

How long readiness has before the pod is failed. Omit to take the source default: 300s for a container app, 2400s for a code app (the probe budget a code worker already had). A later container.yaml deploy overwrites the live value with that document's timeouts.startupSeconds.

appSource

objectrequired

Write-only. Source for the app's first version. Not returned in the App response. Use the /builds endpoints to inspect build status. - code: the codebase is submitted to the build pipeline. A new image is built and deployed once ready. - container: the submitted zip is built into a wrapper image by the same pipeline and deployed once ready. The version records the built image, not a customer-supplied reference. Either way the version is recorded with the app, and the app transitions to active once its first version is ready, or to failed if the build, validation, or rollout fails.

Properties2 properties

type

stringrequired

Selects the version creation path. code submits customer source code to the Image Build Service. container submits a wrapper Dockerfile and container.yaml for Runware to build into a hosted image.

Allowed values2 values

source

variantrequired
Format 1: object3 properties
baseImage
stringrequired

Base image of the served image, e.g. python:3.12-slim. It must provide python 3.12 or newer on its PATH, and the build fails if it does not. The build imports the model file under Python 3.12. If the .python-version or the requires-python of pyproject.toml in the codebase excludes 3.12, the build uses a version that it permits. The Python of the base image must then also be one that they permit, and a .python-version sets the oldest one.

requirements
string[]

Additional pip packages to install alongside the codebase.

codebase
objectrequired
Properties2 properties
sourceId
string (uuid)requiredUUID v4

Id of a source published by a completed upload in this organization (POST /v1/source-uploads, then complete). The archive it names is the zip of the customer's code. A source is immutable and reusable: the same sourceId may back any number of apps and versions, each choosing its own modelFile.

modelFile
stringrequiredmin: 1max: 512

Path of the MLflow model entry point inside the archive, relative to its root (e.g. model.py). Version execution configuration rather than archive identity, so two apps may run one source with different entry points. The build proves the file is in the archive and answers 422 when it is not.

Format 2: object1 property
sourceId
string (uuid)requiredUUID v4

Id of a source published by a completed upload in this organization (POST /v1/source-uploads, then complete). The archive it names carries a wrapper Dockerfile and a container.yaml config document at its root, plus any build-context files the Dockerfile copies in. Runware builds the image from it, resolving the Dockerfile's public base images to immutable digests, and hosts the result, so no image reference or pull credential is supplied: a private base image is not supported until a build-time credential mechanism exists. An invalid container.yaml rejects the create: 400 where the document could not be parsed, 422 where it parsed and broke a rule, with errors[] entries carrying configPointer into the document. The endpoint set it declares becomes visible once the first build is ready and deployed. A source is immutable and reusable: the same sourceId may back any number of apps and versions in the organization.

secrets

object[]

Existing organization secrets to attach to this app, with an optional env-var name override per entry. This is the app's initial attachment set, so the first rollout carries the values into the worker. The secret must already exist and be active. This route does not create one, and an unknown or inactive name returns 404. Shape matches POST /apps/{appId}/secrets so create and attach share one contract. Each injected name must not collide with a key in environmentVariables. See SecretAttach. Each secret can also be attached to at most 25 deployments total, shared with every other route that attaches it. An entry that would push a secret past that returns 422.

Array items2 properties each

secretName

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

envVarName

stringnullablemin: 1max: 128

Environment variable name the secret is injected as. Omit or null to use secretName. The server resolves and stores the final name. Same reserved-name rules as SecretName (422 if reserved). The resolved name must also not collide with a plain environment variable key on this app (422).

Map with environment variables. Keys are the environment variable names, values are the environment variable values. Use the dedicated /environment-variables endpoints to change them after the app exists. Each key must satisfy EnvironmentVariableName. POSIX-style, at most 128 characters. OpenAPI 3.0 cannot constrain map keys, so a bad one is rejected by the server rather than by the schema. Keys must also not collide with a secret's injected env var name on the same app (see attachAppSecret).

volumes

object[]max items: 30

Persistent node-local directories bind-mounted through the checkpointer into the sandboxed application. Use these for downloaded weights and caches that must stay outside the checkpointed root filesystem. Paths must be unique and non-overlapping. The set is frozen into each immutable app version.

Array items1 property each

mountPath

stringrequiredmin: 2max: 2048

Absolute path exposed inside the sandboxed application. The path is the volume's stable identity and is mirrored below the app's node-local data directory. Root, duplicate, and overlapping paths are rejected. Each path component (the text between / separators) is at most 255 bytes, the filesystem NAME_MAX the mirrored node directory must satisfy.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
400The request could not be parsed (bad-request), or it parsed and the container.yaml inside its container source could not be (container-config-yaml-invalid). The second carries errors[]. Where the parser can name a field, an entry has a configPointer into container.yaml and a pointer to the archive in the request body. A document that parses and then breaks a rule returns 422 instead.
401Missing or invalid credentials
402The organization's credit cannot cover the capacity the request asks for. The problem type is insufficient-credit when the available balance is short, and credit-suspended when a refund took back credit already spent and every allocation is refused until the balance is funded back. Either way shortfall is the amount to add before retrying: the request is unchanged by the refusal and succeeds as sent once the credit is there.
404Resource not found
409Resource already exists or the request conflicts with its current state
413The request body exceeds its size limit: 10 MiB on invoke-sync and invoke-async, whose body carries the endpoint's payload, and 1 MiB on the other operations that answer with this response. detail names the limit in bytes.
422The request was well-formed but invalid. A field of the request body that broke the schema or a rule is unprocessable-entity. A container source whose archive or container.yaml broke a contract rule carries the rule's own type and errors[]. An entry about a field of container.yaml has a configPointer to that field and a pointer to the archive in the request body. An entry about the archive has neither. A code source's archive rejection is unprocessable-entity. A request that could not be parsed returns 400.
500Unexpected server error
503A required service is temporarily unavailable

Get an app

GETapi.serverless.runware.ai/v1/apps/{appId}

Returns the app the authenticated organization owns under this appId. An unknown app and a soft-deleted one both return 404 Not Found: a deleted app is gone to its owner, and its rows are retained only for billing and audit. To read deleted apps, list them with status=deleted.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

Update an app

PATCHapi.serverless.runware.ai/v1/apps/{appId}

Patches one or more aspects of an app in place. All fields are optional. Omitted fields are left unchanged. Valid in any non-deleted status, including stopped (changes apply on resume). Lifecycle transitions use the dedicated deploy, stop, resume, and delete operations. A configuration or environmentVariables change records a new version with the same image. If that image is deployable, the update pins it as activeVersionId and rolls the workload when the app is active or initializing. A failed app is moved to initializing and rolled, the same as POST /deploy. If the image is not deployable, the version is recorded and activeVersionId is left unchanged. If the roll fails, activeVersionId is restored and the previous configuration keeps serving. A name-only change records a version and does not pin. A stopped or stopping app pins the version and rolls it on resume. A configuration, environmentVariables, or appSource change while a create or resume rollout is already in progress returns 409 Conflict, as does one to a failed app while the workers of its failed rollout are still stopping. A name-only or secrets-only change does not. appSource starts a build and records version N+1 with a new image tag. The deploy queue carries the build-then-deploy tail. activeVersionId moves only when that rollout completes. A builder rejection (400 where a container document's parser refused it, 422 where it parsed and broke a rule) leaves the app on its current version and writes no version row and no build row. After the builder accepts, version N+1 is recorded even if a concurrent secret deactivation, an env/secret collision, or the 25-deployment attachment ceiling below prevents this request's env/secrets overlay. In that case the previous environmentVariables and attachment set stay in place and are what the new version snapshots, and the 422 the standalone case below returns does not apply here. The request still succeeds. environmentVariables replaces the whole set: a key absent from the map is deleted, and a null value omits that key from the new set. The resolved map is snapshotted onto the new version. secrets replaces the whole attachment set. An attachment absent from the array is detached. Injected names must not collide with a plain environment variable on the app. The combined set of plain variables and attachments is capped at 100, and each individual secret can be attached to at most 25 deployments total. An entry that would push a secret past that returns 422. This is a control-plane record only. Secret values do not reach a pod, and the version snapshot carries no secrets, so a secrets-only change does not roll the workload. Endpoints are not a field of this contract: the set belongs to the app source, so it changes only when a new version with a new source builds and deploys.

An update records a version even when it changes nothing but configuration. That version carries the previous image forward, so a scaling change costs a version and not a build.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Body

appName

stringmin: 1

Mutable display name. Does not affect app identity or routing. Omit to leave unchanged. An explicit blank or padded value is rejected, not treated as a clear.

Partial worker configuration. Any field present overwrites the live value. Omitted fields are left unchanged. Clearing a nullable live field (setting it to null) is not supported. Omit the field to leave it unchanged. fallbackGpuType is the exception: send an empty string to clear it. computeType is create-time only and cannot be patched. Changing gpuType affects only newly created workers.

Properties11 properties

gpuType

stringmin: 1max: 64

Preferred GPU type. Omit to leave unchanged. Rejected with a 422 when no capacity is currently offered for the type (it does not appear in GET /v1/gpu-types).

fallbackGpuType

stringmax: 64

Secondary GPU type. Omit to leave unchanged. Send an empty string to clear. Forms bind an unselected dropdown as "". JSON null is rejected: this field is not nullable, so a client that meant to clear must send "" rather than null. A non-empty value must be an active catalog code. Unlike gpuType this is an existence check only: it does not require admitted capacity, so the code may not appear in the customer GET /v1/gpu-types list.

gpusPerWorker

integerint32

GPUs granted to one worker pod. A worker holds its GPUs as one group the cluster grants indivisibly, so the count is one of the advertised group sizes rather than any number in a range. The value must also be a group size admitted by the cluster backing the chosen gpuType: a count above 1 that cluster does not grant is rejected with a 422 naming /configuration/gpusPerWorker, since the pod could never be scheduled. A value above 1 requires an image built after the multi-GPU worker entrypoint. Older images serve a single rank while holding every granted GPU.

Allowed values4 values

minWorkers

integerint32min: 0

maxWorkers

integerint32min: 1

minAvailableWorkers

integerint32min: 0

Idle workers held above current demand. Omit to leave unchanged. Send 0 to remove the buffer. Null is refused, because omitting a field and clearing it mean different things here. Must be lower than the resolved maxWorkers.

availableWorkersPct

integerint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Omit to leave unchanged. Send 0 to remove the buffer. Null is refused, because omitting a field and clearing it mean different things here.

idleTtlSecs

integerint32min: 1

scalingDelaySecs

integerint32min: 1

requestTimeoutSecs

integerint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Omit to leave unchanged. Applied to container workers only. A code app stores the new value but keeps the platform MLflow read ceiling.

startupTimeoutSecs

integerint32min: 1max: 2400

appSource

object

Write-only. New source to build and deploy. Not returned in the App response. Use the /builds endpoints to inspect build status. Triggers a build (for code sources) or validates container.yaml then builds (for container sources). On accept the resulting version is recorded with a new image tag and rolled through the deploy queue. A builder rejection leaves the app on the previous version and writes no version or build row. After accept, a concurrent secret deactivation or env/secret collision leaves version N+1 recorded and the previous attachment set in place. activeVersionId moves only when that rollout completes. Not valid on a stopped or stopping app: there is nothing to roll the new version onto, and resume rolls the pinned one, so supplying it in those statuses returns 409 Conflict.

Properties2 properties

type

stringrequired

Selects the version creation path. code submits customer source code to the Image Build Service. container submits a wrapper Dockerfile and container.yaml for Runware to build into a hosted image.

Allowed values2 values

source

variantrequired
Format 1: object3 properties
baseImage
stringrequired

Base image of the served image, e.g. python:3.12-slim. It must provide python 3.12 or newer on its PATH, and the build fails if it does not. The build imports the model file under Python 3.12. If the .python-version or the requires-python of pyproject.toml in the codebase excludes 3.12, the build uses a version that it permits. The Python of the base image must then also be one that they permit, and a .python-version sets the oldest one.

requirements
string[]

Additional pip packages to install alongside the codebase.

codebase
objectrequired
Properties2 properties
sourceId
string (uuid)requiredUUID v4

Id of a source published by a completed upload in this organization (POST /v1/source-uploads, then complete). The archive it names is the zip of the customer's code. A source is immutable and reusable: the same sourceId may back any number of apps and versions, each choosing its own modelFile.

modelFile
stringrequiredmin: 1max: 512

Path of the MLflow model entry point inside the archive, relative to its root (e.g. model.py). Version execution configuration rather than archive identity, so two apps may run one source with different entry points. The build proves the file is in the archive and answers 422 when it is not.

Format 2: object1 property
sourceId
string (uuid)requiredUUID v4

Id of a source published by a completed upload in this organization (POST /v1/source-uploads, then complete). The archive it names carries a wrapper Dockerfile and a container.yaml config document at its root, plus any build-context files the Dockerfile copies in. Runware builds the image from it, resolving the Dockerfile's public base images to immutable digests, and hosts the result, so no image reference or pull credential is supplied: a private base image is not supported until a build-time credential mechanism exists. An invalid container.yaml rejects the create: 400 where the document could not be parsed, 422 where it parsed and broke a rule, with errors[] entries carrying configPointer into the document. The endpoint set it declares becomes visible once the first build is ready and deployed. A source is immutable and reusable: the same sourceId may back any number of apps and versions in the organization.

secrets

object[]max items: 100

Replaces the app's secret attachments. Same SecretAttach shape as create and POST /apps/{appId}/secrets. An attachment absent from the array is detached. Injected names must not collide with a plain environment variable on the app. Control-plane record only. Secret values do not reach a pod, and the version snapshot carries no secrets, so this field does not roll the workload. An app holds at most 100 environment bindings in total. This array cannot exceed that ceiling on its own. Each individual secret can also be attached to at most 25 deployments total, shared with every other route that attaches it.

Array items2 properties each

secretName

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

envVarName

stringnullablemin: 1max: 128

Environment variable name the secret is injected as. Omit or null to use secretName. The server resolves and stores the final name. Same reserved-name rules as SecretName (422 if reserved). The resolved name must also not collide with a plain environment variable key on this app (422).

Replaces the app's environment variables. Keys are the variable names, values are the values. A key absent from the map is deleted. A null value omits that key from the new set. The resolved map is snapshotted onto the version this update records, so a deploy applies it. When the copied image is deployable the update pins and rolls, the same as a configuration change.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
400The request could not be parsed (bad-request), or it parsed and the container.yaml inside its container source could not be (container-config-yaml-invalid). The second carries errors[]. Where the parser can name a field, an entry has a configPointer into container.yaml and a pointer to the archive in the request body. A document that parses and then breaks a rule returns 422 instead.
401Missing or invalid credentials
402The organization's credit cannot cover the capacity the request asks for. The problem type is insufficient-credit when the available balance is short, and credit-suspended when a refund took back credit already spent and every allocation is refused until the balance is funded back. Either way shortfall is the amount to add before retrying: the request is unchanged by the refusal and succeeds as sent once the credit is there.
404Resource not found
409Resource already exists or the request conflicts with its current state
413The request body exceeds its size limit: 10 MiB on invoke-sync and invoke-async, whose body carries the endpoint's payload, and 1 MiB on the other operations that answer with this response. detail names the limit in bytes.
422The request was well-formed but invalid. A field of the request body that broke the schema or a rule is unprocessable-entity. A container source whose archive or container.yaml broke a contract rule carries the rule's own type and errors[]. An entry about a field of container.yaml has a configPointer to that field and a pointer to the archive in the request body. An entry about the archive has neither. A code source's archive rejection is unprocessable-entity. A request that could not be parsed returns 400.
500Unexpected server error
503A required service is temporarily unavailable

Delete an app

DELETEapi.serverless.runware.ai/v1/apps/{appId}

Soft delete. Sets status = deleting and returns 202 once that intent is persisted. Router removal, canceling in-progress builds, and worker drain (draining → stopping → stopped) are performed asynchronously by the deployer/Scaler. status becomes deleted once all workers stop. All rows are retained for billing finalization, audit, and usage history. Idempotent if the app is already deleting. The appId is released once status reaches deleted, and not before: while the app is deleting its workload is still being torn down and the name stays taken. A new app created under a released name is a new app and inherits nothing, no version, no build, no event history, and no workers.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

Deploy a version

POSTapi.serverless.runware.ai/v1/apps/{appId}/deploy

Activates a ready version by number, setting activeVersionId and returning 202 once that intent is persisted. Worker rollout, routing switch, and canceling in-progress builds (superseded) are performed asynchronously by the deployer/Scaler. Permitted in any addressable status, including initializing and failed. To roll back, supply an older versionNumber. The operation is identical to a forward deploy. No new version is created and no rebuild happens: the version's existing image is re-applied. Re-deploying the currently active version is permitted and re-applies it. A deploy to a stopped or stopping app records the version and rolls no workload, because no workers are running: the 202 does not imply a rollout there. The recorded version is the one applied when the app resumes. If the roll of a live app fails, activeVersionId is restored to the version that kept serving, so the field keeps naming the running image. Rollout (deployer/Scaler): the platform starts workers on the target version, waits for at least one to become healthy, switches task routing to the new version, then drains old-version workers gracefully. Old workers are given a fixed, platform-managed grace period to finish in-flight tasks before being force-terminated. If new workers fail to become healthy, old workers are not drained and the app continues on the previous version. Readiness wait: a version with minWorkers above zero becomes active only when that many of its workers can serve. The wait starts when the workload is applied, not when this request arrives. It lasts 10 minutes by default. It continues after that, up to 1 hour, while enough of its workers are still pulling or loading their image to reach minWorkers. A first pull of a large image on a node can take this long. If the wait ends first, the rollout fails. A first rollout then moves the app to failed and the platform stops its workers. To recover, deploy a ready version again with this operation once those workers have stopped. Errors: - Deploy to a deleting app returns 409 Conflict - versionNumber not found or not ready returns 409 Conflict - A deploy while a rollout of this app is still in flight returns 409 Conflict. One app rolls to one version at a time. Retry once it completes. - A deploy to a failed app while the workers of its failed rollout are still stopping returns 409 Conflict. Retry once they have stopped. - Deploy to a non-existent or deleted app returns 404 Not Found

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Body

versionNumber

integerrequiredint32

Version number to deploy. Must reference a ready version on this app. Any in-progress or queued builds are canceled. Use an older version number to roll back.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
400The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
401Missing or invalid credentials
402The organization's credit cannot cover the capacity the request asks for. The problem type is insufficient-credit when the available balance is short, and credit-suspended when a refund took back credit already spent and every allocation is refused until the balance is funded back. Either way shortfall is the amount to add before retrying: the request is unchanged by the refusal and succeeds as sent once the credit is there.
404Resource not found
409Resource already exists or the request conflicts with its current state
413The request body exceeds its size limit: 10 MiB on invoke-sync and invoke-async, whose body carries the endpoint's payload, and 1 MiB on the other operations that answer with this response. detail names the limit in bytes.
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

Stop an app

POSTapi.serverless.runware.ai/v1/apps/{appId}/stop

Moves the app to stopping and returns 202 once that intent is persisted. Scale-to-zero and worker drain are performed asynchronously by the Scaler. status becomes stopped once all workers drain. In-flight tasks have a fixed, platform-managed grace period to complete. Workers that exceed it are force-terminated and their tasks return to the queue per delivery guarantees. New task submissions remain accepted while stopping. After the app reaches stopped, submissions return 409 Conflict. Precondition: status is active or failed. The platform already stops the workers of a failed app's rollout. Stopping the app takes it to stopped once they have.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
401Missing or invalid credentials
404Resource not found
409Resource already exists or the request conflicts with its current state
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

Resume an app

POSTapi.serverless.runware.ai/v1/apps/{appId}/resume

Moves the app to initializing and returns 202 once that intent is persisted. The Scaler then starts workers and sets the app to active. Tasks that remained queued when the app stopped are consumed as workers come online. Precondition: status = stopped.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
401Missing or invalid credentials
402The organization's credit cannot cover the capacity the request asks for. The problem type is insufficient-credit when the available balance is short, and credit-suspended when a refund took back credit already spent and every allocation is refused until the balance is funded back. Either way shortfall is the amount to add before retrying: the request is unchanged by the refusal and succeeds as sent once the credit is there.
404Resource not found
409Resource already exists or the request conflicts with its current state
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

Favorite an app

PUTapi.serverless.runware.ai/v1/apps/{appId}/favorite

Pins the app as a favorite for the authenticated organization so the console can surface it in a Favorites section. Idempotent: favoriting an already-favorited app succeeds and returns the app with isFavorite: true. Apps in deleting or deleted status cannot be favorited (404). Soft-delete clears any existing pin when status becomes deleting.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

Remove a favorite

DELETEapi.serverless.runware.ai/v1/apps/{appId}/favorite

Removes the organization favorite pin from the app. Idempotent: unfavoriting an app that is not favorited succeeds and returns the app with isFavorite: false. Valid in any status including deleting and deleted. Unpinning is not an app lifecycle mutation. Missing apps return 404.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Response

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

appName

stringrequiredmin: 1

Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what sort=name orders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.

configuration

objectrequired

Live worker configuration. Updated via PATCH /apps/{appId}.

Properties16 properties

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

computeType

stringrequired

Worker compute class. GPU is the only supported value. CPU workloads are not supported.

Possible values1 value

gpuType

stringnullablemin: 1max: 64

Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.

fallbackGpuType

stringnullablemin: 1max: 64

Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get gpuType waits for that type rather than starting on this one. Do not rely on it as failover.

gpusPerWorker

integerrequiredint32default: 1

GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.

minWorkers

integerrequiredint32min: 0default: 0

Floor for scale-down. 0 = scale to zero.

maxWorkers

integerrequiredint32min: 1

minAvailableWorkers

integernullableint32min: 0

Idle workers held above current demand, so a burst does not wait for a cold start. Null or 0 means no buffer. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers. A buffer below about a tenth of current demand is not added while demand holds steady. availableWorkersPct is not subject to that.

availableWorkersPct

integernullableint32min: 0max: 100

Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to minWorkers.

idleTtlSecs

integerrequiredint32

Seconds a worker can sit idle before the Scaler removes it.

scalingDelaySecs

integerrequiredint32

Cooldown between consecutive scaling decisions.

requestTimeoutSecs

integerrequiredint32min: 1max: 900

How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait: invoke-sync still answers 504 at the platform deadline so the caller can poll. A later container.yaml deploy overwrites this with that document's timeouts.requestSeconds.

startupTimeoutSecs

integerrequiredint32min: 1max: 2400

How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later container.yaml deploy overwrites this with that document's timeouts.startupSeconds.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

health

objectread-only

Whether the app's workload can serve, and why. reason covers more than capacity: a workload that is missing or being torn down also reports state: unavailable. Only capacity_below_floor and demand_unserved are the platform being short of workers, and only those refuse an invocation with capacity-unavailable. The others are refused with the plain service-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only: GET /v1/apps/{appId} and GET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such as stopped, failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.

Properties4 properties

state

stringrequired

Whether the app's workload can serve. healthy can serve. degraded can serve with less capacity than it asks for. unavailable has nothing able to serve. Read AppHealthReason for the cause. This is a report on the workload, not an admission rule. Only an active app is refused on it: an initializing app that already has a version to route to accepts invocations while it reports unavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.

Possible values3 values

reason

stringrequired

Why the app holds its current health state. capacity_below_floor and demand_unserved are the two shapes of capacity exhaustion, and they differ in what the app asked for. capacity_below_floor means the app keeps a warm floor above zero and has fewer workers able to serve than that floor. demand_unserved means the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reports capacity_below_floor from the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Use since to tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart. autoscaler_unhealthy means the autoscaler cannot act on the workload, so the app will not grow with demand. workload_present accompanies a healthy app. workload_terminating and workload_missing are a workload being removed or already gone, which a stop or a delete explains.

Possible values6 values

since

string (date-time)date-time

When the app entered this state. It moves only when state changes, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.

observedAt

string (date-time)date-time

When the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.

effectiveMaxWorkers

integernullableread-onlyint32min: 0

The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.

It describes what was applied, not what is configured now, and the two can differ. It is taken from the maxWorkers of the version that was deployed, so deploying an older version applies that version's ceiling, and a later PATCH of configuration.maxWorkers does not change it until the next deploy. Read it beside configuration.maxWorkers rather than as a bound on it.

null means nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.

runtime

objectrequired

Observed state for one app at calculatedAt. Desired worker scale remains in configuration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.

Properties7 properties

activeWorkers

integerrequiredint64min: 0

Non-terminal workers (status other than stopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.

provisionedGpuCount

integerrequiredint64min: 0

Sum of gpuCount across those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.

calculatedAt

string (date-time)requireddate-time

When this runtime snapshot was read from the database.

requests24h

integerint64min: 0

Requests served by this app in the last 24 hours. Omitted until available.

errorRate24h

numberdoublemin: 0max: 1

Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.

averageRequestDuration24h

numberdoublemin: 0

Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.

queueDepth

integerint64min: 0

Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.

secrets

object[]required

Secrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use /apps/{appId}/secrets to page the set.

Array items7 properties each

id

string (uuid)requiredUUID v4

name

stringrequiredmin: 1max: 128

Organization-scoped secret name. The shape matches EnvironmentVariableName and the secrets.name / deployment_secrets.env_var_name column CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME, DISABLE_NGINX, MLFLOW_MODELS_WORKERS, UVICORN_HOST) are rejected with 422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).

type

stringrequired

Kind of secret. Only the environment-variable variant (generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.

Possible values1 value

metadata

objectnullable

Optional opaque metadata associated with the secret.

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

envVarName

stringnullablemin: 1max: 128

Resolved environment variable name when it differs from name. Omitted when the secret name is used. Same reserved-name rules as SecretName when set.

environmentVariables

object[]required

Plain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the /environment-variables endpoints to page the set.

Array items6 properties each

id

string (uuid)UUID v4

appId

stringmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

key

stringrequiredmin: 1max: 128

POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the deployment_configs column CHECK so a valid-by-contract request cannot 500 at INSERT.

value

stringrequired

createdAt

string (date-time)date-time

updatedAt

string (date-time)date-time

status

stringrequired

Where the app is in its lifecycle. initializing is building or rolling out its first version. active is deployed and accepting invocations. stopping is draining its workers. stopped holds no workers and accepts none. deleting and deleted are removal. failed is a rollout the platform could not complete. A failed app accepts no invocations. Recover it with POST /v1/apps/{appId}/deploy when it has a ready version, or with a new appSource through PATCH /v1/apps/{appId}. Lifecycle is not health. active says the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place stays active. Read health for that.

Possible values7 values

isFavorite

booleanrequired

Whether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via PUT/DELETE /v1/apps/{appId}/favorite.

activeVersionId

string (uuid)nullableUUID v4

Current deployed version. Null until the first version is successfully deployed.

createdAt

string (date-time)requireddate-time

updatedAt

string (date-time)requireddate-time

Errors

StatusWhen
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

Summary across apps

GETapi.serverless.runware.ai/v1/app-summary

Aggregate dashboard metrics across all apps owned by the authenticated organization. App and worker tallies are always present. Request and error-rate totals come from the metrics store and are omitted when that hop cannot answer rather than reported as zero. Spend covers a rolling 24 hours and is omitted when the usage cannot be priced.

Response

activeApps

integerrequiredint64

Apps currently in the active status.

totalApps

integerrequiredint64

All apps in the organization excluding soft-deleted ones.

activeWorkers

integerrequiredint64

Non-terminal workers (status other than stopped) across every app and version in the organization.

provisionedGpuCount

integerrequiredint64

Sum of gpuCount across those same non-terminal workers.

requests24h

integerint64

Requests served in the last 24 hours across every app. Omitted when metrics cannot be read. Zero when the organization had no traffic.

errorRate24h

numberdouble

Error ratio (4xx + 5xx over requests) for the last 24 hours, in 0–1. Omitted when metrics cannot be read or when requests24h is zero.

Provisional pay-as-you-go accrual over the last 24 hours, across every app. Excludes finalization rounding and ledger adjustments. A rolling window rather than a calendar day, so the figure does not reset at a boundary the customer did not choose. What has accrued, not a projection: an app that has just started shows what it has run so far. Time a capacity commitment covered is excluded, because it collects nothing. Omitted when the usage cannot be priced, which a zero would misreport as having run for free.

Properties2 properties

amount

stringrequired

Amount in major units as an exact decimal string.

currency

stringrequired

ISO 4217 alphabetic code. The platform bills in USD only.

Possible values1 value

calculatedAt

string (date-time)requireddate-time

When these metrics were computed.

Errors

StatusWhen
401Missing or invalid credentials
500Unexpected server error
503A required service is temporarily unavailable