Apps
Create an app, change its configuration, deploy a version, and control its lifecycle.
Introduction
An app is the unit you create, configure and invoke. Everything else on this API hangs beneath one.
Creating one names the source it runs, so the archive goes up before the app exists. Source uploads covers that flow and the sourceId it produces.
appId is chosen by you at creation and cannot be changed afterwards. It is released again once the app it named reaches deleted, so a name can be reused by a new app that shares nothing with the old one.
List apps
Returns a page of the organization's apps. Filters combine with AND. Soft-deleted apps are excluded unless status=deleted is requested explicitly. Favorited apps appear before non-favorited apps, with the selected ordering applied within each group.
A cursor is only valid for the sort and filters it was issued under. Reusing one across a different ordering or filter set returns 400.
Request
limit
integerint32min: 1max: 100default: 20Maximum number of items to return.
cursor
stringOpaque pagination cursor returned by a previous call, as
nextCursoror, on the operations that offer one,prevCursor.
status
stringReturn only apps in this status.
Allowed values7 values
q
stringmin: 1max: 100Case-insensitive substring match against
appNameandappId. An app matching either is returned.
gpuType
stringmin: 1max: 64Return only apps whose worker configuration requests this GPU type. Matched against
configuration.gpuTypeonly, notfallbackGpuType. Must be a code fromGET /v1/gpu-types. An unknown code returns422.
sort
stringdefault: createdAtOrdering for
listApps. Favorited apps appear before non-favorited apps, and the selected ordering applies within each group. Every ordering is total (ties broken byappId), so a page is reproducible and its cursor stable. -createdAt: newest first. The default. -name:appNameA–Z, case-insensitive.Allowed values2 values
Response
nextCursor
stringnullableCursor for the next page. Null when there are no more items.
data
object[]Array items13 properties each
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
400 | The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
|
401 | Missing or invalid credentials |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Create an app
Creates an app together with its worker configuration, environment variables and endpoints, and records version 1, the immutable description of what to deploy. The app starts in initializing, and what happens next depends on the app source type:
- code source: the codebase is submitted to the build pipeline. Once the image is built
and workers become healthy the app transitions to active and activeVersionId
points at that version. If the build, validation, or rollout fails the app is
marked failed.
- container source: the submitted zip (wrapper Dockerfile + container.yaml)
goes through the same build pipeline. The wrapper image is built, published and
deployed, so the version carries a buildId and the app follows the same
lifecycle as a code source. An invalid container.yaml rejects the create
before any build capacity is spent: 400 where the document could not be
parsed at all, 422 where it parsed and broke a rule.
activeVersionId is null until a rollout completes: a version records what should run, and only a finished deploy says what does.
secrets attaches organization secrets that already exist. It is the app's initial attachment set, so the first rollout carries their values into the worker. This route does not create a secret. Use POST /v1/secrets first. A name that is unknown to the organization, or that is not active, returns 404. A name that collides with a key in environmentVariables, a repeated name and a set that goes past the binding limit each return 422. The whole set is checked before any build capacity is spent.
Creating an app does not make it callable. The first build has to finish and roll out, and until it does the app stays initializing and answers 409.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredProperties12 properties
computeType
stringdefault: gpuGPU is the only supported compute type. Omitting the field selects GPU. CPU workloads are not supported. A request that names
cpuis rejected with 422 before a build or deploy starts.Allowed values1 value
gpuType
stringrequiredmin: 1max: 64GPU type the workers run on. Required: omitting it (or sending null) is a 422 before a build or deploy starts, because a GPU app with no type is unpinned and the deployer would render NVIDIA defaults. Must match an
idreturned byGET /v1/gpu-typesthat currently has admitted capacity.
fallbackGpuType
stringnullablemax: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it, so it is not failover you can rely on today. Omit, send JSON null, or send an empty string for no fallback. Forms bind an unselected dropdown as
"", which is not aGpuTypeId. A non-empty value must be an active catalog code. UnlikegpuTypethis is an existence check only: it does not require admitted capacity, so the code may not appear in the customerGET /v1/gpu-typeslist.
gpusPerWorker
integerint32default: 1GPUs granted to one worker pod. A worker holds its GPUs as one group the cluster grants indivisibly, so the count is one of the advertised group sizes rather than any number in a range. The value must also be a group size admitted by the cluster backing the chosen
gpuType: a count above 1 that cluster does not grant is rejected with a 422 naming/configuration/gpusPerWorker, since the pod could never be scheduled. A value above 1 requires an image built after the multi-GPU worker entrypoint. Older images serve a single rank while holding every granted GPU.Allowed values4 values
minWorkers
integerint32min: 0default: 0
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Omit, or send null or 0, for no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty, so it does not stop an idle app scaling to
minWorkers.
idleTtlSecs
integerrequiredint32min: 1
scalingDelaySecs
integerrequiredint32min: 1
requestTimeoutSecs
integerint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only. Omit on a container app to take 60s. A code app may still send or omit this field: omit stores 600s for the console, and neither value is applied to MLflow reads, which keep the platform ceiling. A later
container.yamldeploy overwrites the live value with that document'stimeouts.requestSeconds. Bounds the worker, not the synchronous HTTP wait.
startupTimeoutSecs
integerint32min: 1max: 2400How long readiness has before the pod is failed. Omit to take the source default: 300s for a container app, 2400s for a code app (the probe budget a code worker already had). A later
container.yamldeploy overwrites the live value with that document'stimeouts.startupSeconds.
appSource
objectrequiredWrite-only. Source for the app's first version. Not returned in the App response. Use the
/buildsendpoints to inspect build status. -code: the codebase is submitted to the build pipeline. A new image is built and deployed once ready. -container: the submitted zip is built into a wrapper image by the same pipeline and deployed once ready. The version records the built image, not a customer-supplied reference. Either way the version is recorded with the app, and the app transitions toactiveonce its first version is ready, or tofailedif the build, validation, or rollout fails.Properties2 properties
type
stringrequiredSelects the version creation path.
codesubmits customer source code to the Image Build Service.containersubmits a wrapperDockerfileandcontainer.yamlfor Runware to build into a hosted image.Allowed values2 values
source
variantrequiredFormat 1: object3 properties
baseImage
stringrequiredBase image of the served image, e.g.
python:3.12-slim. It must providepython3.12 or newer on its PATH, and the build fails if it does not. The build imports the model file under Python 3.12. If the.python-versionor therequires-pythonofpyproject.tomlin the codebase excludes 3.12, the build uses a version that it permits. The Python of the base image must then also be one that they permit, and a.python-versionsets the oldest one.
requirements
string[]Additional pip packages to install alongside the codebase.
codebase
objectrequiredProperties2 properties
sourceId
string (uuid)requiredUUID v4Id of a source published by a completed upload in this organization (
POST /v1/source-uploads, thencomplete). The archive it names is the zip of the customer's code. A source is immutable and reusable: the samesourceIdmay back any number of apps and versions, each choosing its ownmodelFile.
modelFile
stringrequiredmin: 1max: 512Path of the MLflow model entry point inside the archive, relative to its root (e.g.
model.py). Version execution configuration rather than archive identity, so two apps may run one source with different entry points. The build proves the file is in the archive and answers422when it is not.
Format 2: object1 property
sourceId
string (uuid)requiredUUID v4Id of a source published by a completed upload in this organization (
POST /v1/source-uploads, thencomplete). The archive it names carries a wrapperDockerfileand acontainer.yamlconfig document at its root, plus any build-context files the Dockerfile copies in. Runware builds the image from it, resolving the Dockerfile's public base images to immutable digests, and hosts the result, so no image reference or pull credential is supplied: a private base image is not supported until a build-time credential mechanism exists. An invalidcontainer.yamlrejects the create:400where the document could not be parsed,422where it parsed and broke a rule, witherrors[]entries carryingconfigPointerinto the document. The endpoint set it declares becomes visible once the first build is ready and deployed. A source is immutable and reusable: the samesourceIdmay back any number of apps and versions in the organization.
secrets
object[]Existing organization secrets to attach to this app, with an optional env-var name override per entry. This is the app's initial attachment set, so the first rollout carries the values into the worker. The secret must already exist and be
active. This route does not create one, and an unknown or inactive name returns404. Shape matchesPOST /apps/{appId}/secretsso create and attach share one contract. Each injected name must not collide with a key inenvironmentVariables. SeeSecretAttach. Each secret can also be attached to at most 25 deployments total, shared with every other route that attaches it. An entry that would push a secret past that returns422.Array items2 properties each
secretName
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
envVarName
stringnullablemin: 1max: 128Environment variable name the secret is injected as. Omit or null to use
secretName. The server resolves and stores the final name. Same reserved-name rules asSecretName(422if reserved). The resolved name must also not collide with a plain environment variable key on this app (422).
environmentVariables
objectMap with environment variables. Keys are the environment variable names, values are the environment variable values. Use the dedicated
/environment-variablesendpoints to change them after the app exists. Each key must satisfyEnvironmentVariableName. POSIX-style, at most 128 characters. OpenAPI 3.0 cannot constrain map keys, so a bad one is rejected by the server rather than by the schema. Keys must also not collide with a secret's injected env var name on the same app (seeattachAppSecret).
volumes
object[]max items: 30Persistent node-local directories bind-mounted through the checkpointer into the sandboxed application. Use these for downloaded weights and caches that must stay outside the checkpointed root filesystem. Paths must be unique and non-overlapping. The set is frozen into each immutable app version.
Array items1 property each
mountPath
stringrequiredmin: 2max: 2048Absolute path exposed inside the sandboxed application. The path is the volume's stable identity and is mirrored below the app's node-local data directory. Root, duplicate, and overlapping paths are rejected. Each path component (the text between
/separators) is at most 255 bytes, the filesystem NAME_MAX the mirrored node directory must satisfy.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
400 | The request could not be parsed (bad-request), or it parsed and the container.yaml inside its container source could not be (container-config-yaml-invalid). The second carries errors[]. Where the parser can name a field, an entry has a configPointer into container.yaml and a pointer to the archive in the request body. A document that parses and then breaks a rule returns 422 instead.
|
401 | Missing or invalid credentials |
402 | The organization's credit cannot cover the capacity the request asks for. The problem type is insufficient-credit when the available balance is short, and credit-suspended when a refund took back credit already spent and every allocation is refused until the balance is funded back. Either way shortfall is the amount to add before retrying: the request is unchanged by the refusal and succeeds as sent once the credit is there.
|
404 | Resource not found |
409 | Resource already exists or the request conflicts with its current state |
413 | The request body exceeds its size limit: 10 MiB on invoke-sync and invoke-async, whose body carries the endpoint's payload, and 1 MiB on the other operations that answer with this response. detail names the limit in bytes.
|
422 | The request was well-formed but invalid. A field of the request body that broke the schema or a rule is unprocessable-entity. A container source whose archive or container.yaml broke a contract rule carries the rule's own type and errors[]. An entry about a field of container.yaml has a configPointer to that field and a pointer to the archive in the request body. An entry about the archive has neither. A code source's archive rejection is unprocessable-entity. A request that could not be parsed returns 400.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Get an app
Returns the app the authenticated organization owns under this appId. An unknown app and a soft-deleted one both return 404 Not Found: a deleted app is gone to its owner, and its rows are retained only for billing and audit. To read deleted apps, list them with status=deleted.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Update an app
Patches one or more aspects of an app in place. All fields are optional. Omitted fields are left unchanged. Valid in any non-deleted status, including stopped (changes apply on resume). Lifecycle transitions use the dedicated deploy, stop, resume, and delete operations.
A configuration or environmentVariables change records a new version with the same image. If that image is deployable, the update pins it as activeVersionId and rolls the workload when the app is active or initializing. A failed app is moved to initializing and rolled, the same as POST /deploy. If the image is not deployable, the version is recorded and activeVersionId is left unchanged. If the roll fails, activeVersionId is restored and the previous configuration keeps serving. A name-only change records a version and does not pin. A stopped or stopping app pins the version and rolls it on resume. A configuration, environmentVariables, or appSource change while a create or resume rollout is already in progress returns 409 Conflict, as does one to a failed app while the workers of its failed rollout are still stopping. A name-only or secrets-only change does not.
appSource starts a build and records version N+1 with a new image tag. The deploy queue carries the build-then-deploy tail. activeVersionId moves only when that rollout completes. A builder rejection (400 where a container document's parser refused it, 422 where it parsed and broke a rule) leaves the app on its current version and writes no version row and no build row. After the builder accepts, version N+1 is recorded even if a concurrent secret deactivation, an env/secret collision, or the 25-deployment attachment ceiling below prevents this request's env/secrets overlay. In that case the previous environmentVariables and attachment set stay in place and are what the new version snapshots, and the 422 the standalone case below returns does not apply here. The request still succeeds.
environmentVariables replaces the whole set: a key absent from the map is deleted, and a null value omits that key from the new set. The resolved map is snapshotted onto the new version.
secrets replaces the whole attachment set. An attachment absent from the array is detached. Injected names must not collide with a plain environment variable on the app. The combined set of plain variables and attachments is capped at 100, and each individual secret can be attached to at most 25 deployments total. An entry that would push a secret past that returns 422. This is a control-plane record only. Secret values do not reach a pod, and the version snapshot carries no secrets, so a secrets-only change does not roll the workload.
Endpoints are not a field of this contract: the set belongs to the app source, so it changes only when a new version with a new source builds and deploys.
An update records a version even when it changes nothing but configuration. That version carries the previous image forward, so a scaling change costs a version and not a build.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
appName
stringmin: 1Mutable display name. Does not affect app identity or routing. Omit to leave unchanged. An explicit blank or padded value is rejected, not treated as a clear.
configuration
objectPartial worker configuration. Any field present overwrites the live value. Omitted fields are left unchanged. Clearing a nullable live field (setting it to null) is not supported. Omit the field to leave it unchanged.
fallbackGpuTypeis the exception: send an empty string to clear it.computeTypeis create-time only and cannot be patched. ChanginggpuTypeaffects only newly created workers.Properties11 properties
gpuType
stringmin: 1max: 64Preferred GPU type. Omit to leave unchanged. Rejected with a 422 when no capacity is currently offered for the type (it does not appear in
GET /v1/gpu-types).
fallbackGpuType
stringmax: 64Secondary GPU type. Omit to leave unchanged. Send an empty string to clear. Forms bind an unselected dropdown as
"". JSON null is rejected: this field is not nullable, so a client that meant to clear must send""rather than null. A non-empty value must be an active catalog code. UnlikegpuTypethis is an existence check only: it does not require admitted capacity, so the code may not appear in the customerGET /v1/gpu-typeslist.
gpusPerWorker
integerint32GPUs granted to one worker pod. A worker holds its GPUs as one group the cluster grants indivisibly, so the count is one of the advertised group sizes rather than any number in a range. The value must also be a group size admitted by the cluster backing the chosen
gpuType: a count above 1 that cluster does not grant is rejected with a 422 naming/configuration/gpusPerWorker, since the pod could never be scheduled. A value above 1 requires an image built after the multi-GPU worker entrypoint. Older images serve a single rank while holding every granted GPU.Allowed values4 values
minWorkers
integerint32min: 0
maxWorkers
integerint32min: 1
availableWorkersPct
integerint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Omit to leave unchanged. Send 0 to remove the buffer. Null is refused, because omitting a field and clearing it mean different things here.
idleTtlSecs
integerint32min: 1
scalingDelaySecs
integerint32min: 1
requestTimeoutSecs
integerint32min: 1max: 900How long the sidecar waits on one forwarded container request. Omit to leave unchanged. Applied to container workers only. A code app stores the new value but keeps the platform MLflow read ceiling.
startupTimeoutSecs
integerint32min: 1max: 2400
appSource
objectWrite-only. New source to build and deploy. Not returned in the App response. Use the
/buildsendpoints to inspect build status. Triggers a build (forcodesources) or validatescontainer.yamlthen builds (forcontainersources). On accept the resulting version is recorded with a new image tag and rolled through the deploy queue. A builder rejection leaves the app on the previous version and writes no version or build row. After accept, a concurrent secret deactivation or env/secret collision leaves version N+1 recorded and the previous attachment set in place.activeVersionIdmoves only when that rollout completes. Not valid on astoppedorstoppingapp: there is nothing to roll the new version onto, andresumerolls the pinned one, so supplying it in those statuses returns409 Conflict.Properties2 properties
type
stringrequiredSelects the version creation path.
codesubmits customer source code to the Image Build Service.containersubmits a wrapperDockerfileandcontainer.yamlfor Runware to build into a hosted image.Allowed values2 values
source
variantrequiredFormat 1: object3 properties
baseImage
stringrequiredBase image of the served image, e.g.
python:3.12-slim. It must providepython3.12 or newer on its PATH, and the build fails if it does not. The build imports the model file under Python 3.12. If the.python-versionor therequires-pythonofpyproject.tomlin the codebase excludes 3.12, the build uses a version that it permits. The Python of the base image must then also be one that they permit, and a.python-versionsets the oldest one.
requirements
string[]Additional pip packages to install alongside the codebase.
codebase
objectrequiredProperties2 properties
sourceId
string (uuid)requiredUUID v4Id of a source published by a completed upload in this organization (
POST /v1/source-uploads, thencomplete). The archive it names is the zip of the customer's code. A source is immutable and reusable: the samesourceIdmay back any number of apps and versions, each choosing its ownmodelFile.
modelFile
stringrequiredmin: 1max: 512Path of the MLflow model entry point inside the archive, relative to its root (e.g.
model.py). Version execution configuration rather than archive identity, so two apps may run one source with different entry points. The build proves the file is in the archive and answers422when it is not.
Format 2: object1 property
sourceId
string (uuid)requiredUUID v4Id of a source published by a completed upload in this organization (
POST /v1/source-uploads, thencomplete). The archive it names carries a wrapperDockerfileand acontainer.yamlconfig document at its root, plus any build-context files the Dockerfile copies in. Runware builds the image from it, resolving the Dockerfile's public base images to immutable digests, and hosts the result, so no image reference or pull credential is supplied: a private base image is not supported until a build-time credential mechanism exists. An invalidcontainer.yamlrejects the create:400where the document could not be parsed,422where it parsed and broke a rule, witherrors[]entries carryingconfigPointerinto the document. The endpoint set it declares becomes visible once the first build is ready and deployed. A source is immutable and reusable: the samesourceIdmay back any number of apps and versions in the organization.
secrets
object[]max items: 100Replaces the app's secret attachments. Same
SecretAttachshape as create andPOST /apps/{appId}/secrets. An attachment absent from the array is detached. Injected names must not collide with a plain environment variable on the app. Control-plane record only. Secret values do not reach a pod, and the version snapshot carries no secrets, so this field does not roll the workload. An app holds at most 100 environment bindings in total. This array cannot exceed that ceiling on its own. Each individual secret can also be attached to at most 25 deployments total, shared with every other route that attaches it.Array items2 properties each
secretName
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
envVarName
stringnullablemin: 1max: 128Environment variable name the secret is injected as. Omit or null to use
secretName. The server resolves and stores the final name. Same reserved-name rules asSecretName(422if reserved). The resolved name must also not collide with a plain environment variable key on this app (422).
environmentVariables
objectReplaces the app's environment variables. Keys are the variable names, values are the values. A key absent from the map is deleted. A null value omits that key from the new set. The resolved map is snapshotted onto the version this update records, so a deploy applies it. When the copied image is deployable the update pins and rolls, the same as a configuration change.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
400 | The request could not be parsed (bad-request), or it parsed and the container.yaml inside its container source could not be (container-config-yaml-invalid). The second carries errors[]. Where the parser can name a field, an entry has a configPointer into container.yaml and a pointer to the archive in the request body. A document that parses and then breaks a rule returns 422 instead.
|
401 | Missing or invalid credentials |
402 | The organization's credit cannot cover the capacity the request asks for. The problem type is insufficient-credit when the available balance is short, and credit-suspended when a refund took back credit already spent and every allocation is refused until the balance is funded back. Either way shortfall is the amount to add before retrying: the request is unchanged by the refusal and succeeds as sent once the credit is there.
|
404 | Resource not found |
409 | Resource already exists or the request conflicts with its current state |
413 | The request body exceeds its size limit: 10 MiB on invoke-sync and invoke-async, whose body carries the endpoint's payload, and 1 MiB on the other operations that answer with this response. detail names the limit in bytes.
|
422 | The request was well-formed but invalid. A field of the request body that broke the schema or a rule is unprocessable-entity. A container source whose archive or container.yaml broke a contract rule carries the rule's own type and errors[]. An entry about a field of container.yaml has a configPointer to that field and a pointer to the archive in the request body. An entry about the archive has neither. A code source's archive rejection is unprocessable-entity. A request that could not be parsed returns 400.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Delete an app
Soft delete. Sets status = deleting and returns 202 once that intent is persisted. Router removal, canceling in-progress builds, and worker drain (draining → stopping → stopped) are performed asynchronously by the deployer/Scaler. status becomes deleted once all workers stop. All rows are retained for billing finalization, audit, and usage history. Idempotent if the app is already deleting.
The appId is released once status reaches deleted, and not before: while the app is deleting its workload is still being torn down and the name stays taken. A new app created under a released name is a new app and inherits nothing, no version, no build, no event history, and no workers.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Deploy a version
Activates a ready version by number, setting activeVersionId and returning 202 once that intent is persisted. Worker rollout, routing switch, and canceling in-progress builds (superseded) are performed asynchronously by the deployer/Scaler. Permitted in any addressable status, including initializing and failed.
To roll back, supply an older versionNumber. The operation is identical to a forward deploy. No new version is created and no rebuild happens: the version's existing image is re-applied. Re-deploying the currently active version is permitted and re-applies it.
A deploy to a stopped or stopping app records the version and rolls no workload, because no workers are running: the 202 does not imply a rollout there. The recorded version is the one applied when the app resumes.
If the roll of a live app fails, activeVersionId is restored to the version that kept serving, so the field keeps naming the running image.
Rollout (deployer/Scaler): the platform starts workers on the target version, waits for at least one to become healthy, switches task routing to the new version, then drains old-version workers gracefully. Old workers are given a fixed, platform-managed grace period to finish in-flight tasks before being force-terminated. If new workers fail to become healthy, old workers are not drained and the app continues on the previous version.
Readiness wait: a version with minWorkers above zero becomes active only when that many of its workers can serve. The wait starts when the workload is applied, not when this request arrives. It lasts 10 minutes by default. It continues after that, up to 1 hour, while enough of its workers are still pulling or loading their image to reach minWorkers. A first pull of a large image on a node can take this long. If the wait ends first, the rollout fails. A first rollout then moves the app to failed and the platform stops its workers. To recover, deploy a ready version again with this operation once those workers have stopped.
Errors: - Deploy to a deleting app returns 409 Conflict - versionNumber not found or not ready returns 409 Conflict - A deploy while a rollout of this app is still in flight returns 409 Conflict.
One app rolls to one version at a time. Retry once it completes.
- A deploy to a failed app while the workers of its failed rollout are still
stopping returns 409 Conflict. Retry once they have stopped.
- Deploy to a non-existent or deleted app returns 404 Not Found
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
versionNumber
integerrequiredint32Version number to deploy. Must reference a
readyversion on this app. Any in-progress or queued builds are canceled. Use an older version number to roll back.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
400 | The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
|
401 | Missing or invalid credentials |
402 | The organization's credit cannot cover the capacity the request asks for. The problem type is insufficient-credit when the available balance is short, and credit-suspended when a refund took back credit already spent and every allocation is refused until the balance is funded back. Either way shortfall is the amount to add before retrying: the request is unchanged by the refusal and succeeds as sent once the credit is there.
|
404 | Resource not found |
409 | Resource already exists or the request conflicts with its current state |
413 | The request body exceeds its size limit: 10 MiB on invoke-sync and invoke-async, whose body carries the endpoint's payload, and 1 MiB on the other operations that answer with this response. detail names the limit in bytes.
|
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Stop an app
Moves the app to stopping and returns 202 once that intent is persisted. Scale-to-zero and worker drain are performed asynchronously by the Scaler. status becomes stopped once all workers drain. In-flight tasks have a fixed, platform-managed grace period to complete. Workers that exceed it are force-terminated and their tasks return to the queue per delivery guarantees. New task submissions remain accepted while stopping. After the app reaches stopped, submissions return 409 Conflict. Precondition: status is active or failed. The platform already stops the workers of a failed app's rollout. Stopping the app takes it to stopped once they have.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
401 | Missing or invalid credentials |
404 | Resource not found |
409 | Resource already exists or the request conflicts with its current state |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Resume an app
Moves the app to initializing and returns 202 once that intent is persisted. The Scaler then starts workers and sets the app to active. Tasks that remained queued when the app stopped are consumed as workers come online. Precondition: status = stopped.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
401 | Missing or invalid credentials |
402 | The organization's credit cannot cover the capacity the request asks for. The problem type is insufficient-credit when the available balance is short, and credit-suspended when a refund took back credit already spent and every allocation is refused until the balance is funded back. Either way shortfall is the amount to add before retrying: the request is unchanged by the refusal and succeeds as sent once the credit is there.
|
404 | Resource not found |
409 | Resource already exists or the request conflicts with its current state |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Favorite an app
Pins the app as a favorite for the authenticated organization so the console can surface it in a Favorites section. Idempotent: favoriting an already-favorited app succeeds and returns the app with isFavorite: true. Apps in deleting or deleted status cannot be favorited (404). Soft-delete clears any existing pin when status becomes deleting.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Remove a favorite
Removes the organization favorite pin from the app. Idempotent: unfavoriting an app that is not favorited succeeds and returns the app with isFavorite: false. Valid in any status including deleting and deleted. Unpinning is not an app lifecycle mutation. Missing apps return 404.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
Response
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
appName
stringrequiredmin: 1Mutable display name. Must start and end with a non-whitespace character: it is what the console renders and what
sort=nameorders on, and it is not required to be unique. Interior spaces are allowed ("Sentiment Analysis"). Leading or trailing whitespace is rejected, because a padded name is indistinguishable from its trimmed form in the console and breaks a name-confirm delete.
configuration
objectrequiredLive worker configuration. Updated via
PATCH /apps/{appId}.Properties16 properties
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
computeType
stringrequiredWorker compute class. GPU is the only supported value. CPU workloads are not supported.
Possible values1 value
gpuType
stringnullablemin: 1max: 64Preferred GPU type. Absent (or null) only on historical apps created before a GPU type was required.
fallbackGpuType
stringnullablemin: 1max: 64Secondary GPU type recorded for this app. It is validated and stored, but placement does not yet substitute it: a worker that cannot get
gpuTypewaits for that type rather than starting on this one. Do not rely on it as failover.
gpusPerWorker
integerrequiredint32default: 1GPUs granted to one worker pod. Create and update accept only the group sizes the cluster grants indivisibly, since a worker holds its GPUs as one such group. Historical apps may contain another value.
minWorkers
integerrequiredint32min: 0default: 0Floor for scale-down. 0 = scale to zero.
maxWorkers
integerrequiredint32min: 1
availableWorkersPct
integernullableint32min: 0max: 100Idle workers held above current demand, as a percentage of that demand, rounded up. Null or 0 means no buffer. When both buffers are set the larger of the two applies. The buffer applies only while the queue is non-empty: an idle app still scales to
minWorkers.
idleTtlSecs
integerrequiredint32Seconds a worker can sit idle before the Scaler removes it.
scalingDelaySecs
integerrequiredint32Cooldown between consecutive scaling decisions.
requestTimeoutSecs
integerrequiredint32min: 1max: 900How long the sidecar waits on one forwarded container request. Container-only: a code app keeps the platform MLflow read ceiling, so this field is stored and returned but is not applied to those reads. Bounds the worker, not the synchronous HTTP wait:
invoke-syncstill answers 504 at the platform deadline so the caller can poll. A latercontainer.yamldeploy overwrites this with that document'stimeouts.requestSeconds.
startupTimeoutSecs
integerrequiredint32min: 1max: 2400How long readiness has before the pod is failed. Rendered as the startup probe budget and the sidecar's own startup gate. A later
container.yamldeploy overwrites this with that document'stimeouts.startupSeconds.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
health
objectread-onlyWhether the app's workload can serve, and why.
reasoncovers more than capacity: a workload that is missing or being torn down also reportsstate: unavailable. Onlycapacity_below_flooranddemand_unservedare the platform being short of workers, and only those refuse an invocation withcapacity-unavailable. The others are refused with the plainservice-unavailable, because retrying does not bring a removed workload back. Populated on the app reads only:GET /v1/apps/{appId}andGET /v1/apps. Every other response that carries an app omits it, because the verdict is observed from the cluster rather than changed by the request: read the app again after a mutation. Within those two reads it is absent for any of four reasons: the platform has not observed the app yet, which is normal for one that has never deployed. The app is in a lifecycle state whose verdict is not published, such asstopped,failed, or one that is draining. The verdict could not be read on this request, or the stored verdict carries a value this version of the API does not recognise. The four are not distinguished, so absence is never a claim that the app is healthy. The first two are stable, the third clears by itself, and the last persists until the API is upgraded.Properties4 properties
state
stringrequiredWhether the app's workload can serve.
healthycan serve.degradedcan serve with less capacity than it asks for.unavailablehas nothing able to serve. ReadAppHealthReasonfor the cause. This is a report on the workload, not an admission rule. Only anactiveapp is refused on it: aninitializingapp that already has a version to route to accepts invocations while it reportsunavailable, which is the ordinary case during a first deploy, and a draining app is not gated on health at all. Do not read this field as whether the next invocation will be accepted.Possible values3 values
reason
stringrequiredWhy the app holds its current health state.
capacity_below_flooranddemand_unservedare the two shapes of capacity exhaustion, and they differ in what the app asked for.capacity_below_floormeans the app keeps a warm floor above zero and has fewer workers able to serve than that floor.demand_unservedmeans the app scales to zero, so it has no floor to be short of, and work is waiting on its queue with nothing running it. They behave differently during a cold start. An app waking from zero is given a grace period before it is called starved, so an ordinary wake-up is not reported as a fault. An app with a warm floor gets no such grace: it reportscapacity_below_floorfrom the moment its workload is applied until its first worker is ready, so a normal first deploy reports it for the whole of its cold start. Usesinceto tell the two apart. A cold start clears within the app's startup time, and a real shortage does not. Neither names whose fault the shortfall is, and neither is a statement about charging. A worker the platform never placed holds no GPU and costs nothing, but the same two reasons also cover workers that were placed and cannot serve, and those hold a GPU. Some of those states are charged for and some are not: a container that crash-loops or one still loading is charged, while one still pulling its image is not. Read the app's workers to tell the cases apart.autoscaler_unhealthymeans the autoscaler cannot act on the workload, so the app will not grow with demand.workload_presentaccompanies a healthy app.workload_terminatingandworkload_missingare a workload being removed or already gone, which a stop or a delete explains.Possible values6 values
since
string (date-time)date-timeWhen the app entered this state. It moves only when
statechanges, so it answers how long the condition has held, the figure to quote when asking how long an app has been unable to serve.
observedAt
string (date-time)date-timeWhen the platform last looked. It is rewritten on every observation, so it reports the freshness of the verdict and not the age of the condition. A value far in the past means nothing has observed the app recently.
effectiveMaxWorkers
integernullableread-onlyint32min: 0The worker ceiling the last deploy actually applied, reduced where the organization's credit balance did not back the whole range. The autoscaler cannot grow past it.
It describes what was applied, not what is configured now, and the two can differ. It is taken from the
maxWorkersof the version that was deployed, so deploying an older version applies that version's ceiling, and a laterPATCHofconfiguration.maxWorkersdoes not change it until the next deploy. Read it besideconfiguration.maxWorkersrather than as a bound on it.nullmeans nothing has been applied yet. It is recalculated on every deploy. A credit top-up also recalculates a reduced ceiling and restores the funded range, with no redeploy.
runtime
objectrequiredObserved state for one app at
calculatedAt. Desired worker scale remains inconfiguration. Worker and GPU counts are always present. Traffic, duration, and queue fields are omitted when their backing data is unavailable.Properties7 properties
activeWorkers
integerrequiredint64min: 0Non-terminal workers (
statusother thanstopped) on this app's active version. Pending workers count because they appear in the active-version workers list before Kubernetes assigns a GPU. Outgoing-version workers are excluded because the default workers list omits them. A deleted app reports zero.
provisionedGpuCount
integerrequiredint64min: 0Sum of
gpuCountacross those workers. A pending worker contributes zero until Kubernetes schedules it onto a node.
calculatedAt
string (date-time)requireddate-timeWhen this runtime snapshot was read from the database.
requests24h
integerint64min: 0Requests served by this app in the last 24 hours. Omitted until available.
errorRate24h
numberdoublemin: 0max: 1Error ratio (4xx + 5xx over requests) for this app in the last 24 hours, in 0–1. Omitted when metrics cannot be read or when the app had no requests in the window.
averageRequestDuration24h
numberdoublemin: 0Mean inference request duration in seconds over the last 24 hours (all requests, the duration histogram has no status class). Omitted when metrics cannot be read, when the app had no requests, or when the duration series has no samples for the app in the window.
queueDepth
integerint64min: 0Ready plus unacknowledged messages on this app's live inference queue. Omitted when the app is not live or the gauge cannot be read. Zero when the queue is live and empty.
secrets
object[]requiredSecrets attached to this app, including any env-var name override. Populated on single-app responses. List of apps returns an empty array to avoid an N+1. Use
/apps/{appId}/secretsto page the set.Array items7 properties each
id
string (uuid)requiredUUID v4
name
stringrequiredmin: 1max: 128Organization-scoped secret name. The shape matches
EnvironmentVariableNameand thesecrets.name/deployment_secrets.env_var_namecolumn CHECKs, one rule for the contract and the schema, because attached secrets are intended to be injected as environment variables once ADR-019 in-pod unseal lands. Names the platform sets on the serving container (RUNTIME,DISABLE_NGINX,MLFLOW_MODELS_WORKERS,UVICORN_HOST) are rejected with422: when injection exists, the deployer appends customer env after its own and kubelet resolves duplicates last-wins, so an accepted collision would silently replace a platform value. Same guard as plain environment variables. Enforced by the server (not expressible as a pattern here).
type
stringrequiredKind of secret. Only the environment-variable variant (
generic) is supported. Image-pull (registry) credentials are consumed by the kubelet before any container starts, so they cannot use the in-pod unseal path (ADR-019) and await their own decision.Possible values1 value
metadata
objectnullableOptional opaque metadata associated with the secret.
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
envVarName
stringnullablemin: 1max: 128Resolved environment variable name when it differs from
name. Omitted when the secret name is used. Same reserved-name rules asSecretNamewhen set.
environmentVariables
object[]requiredPlain-text environment variables for this app. Populated on single-app responses (get, update, stop, resume, delete, deploy, favorite). List of apps returns an empty array to avoid an N+1 per page row. Use the
/environment-variablesendpoints to page the set.Array items6 properties each
id
string (uuid)UUID v4
appId
stringmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
key
stringrequiredmin: 1max: 128POSIX-style environment variable name. Letters, digits, and underscore. Must start with a letter or underscore. Matched by the
deployment_configscolumn CHECK so a valid-by-contract request cannot 500 at INSERT.
value
stringrequired
createdAt
string (date-time)date-time
updatedAt
string (date-time)date-time
status
stringrequiredWhere the app is in its lifecycle.
initializingis building or rolling out its first version.activeis deployed and accepting invocations.stoppingis draining its workers.stoppedholds no workers and accepts none.deletinganddeletedare removal.failedis a rollout the platform could not complete. Afailedapp accepts no invocations. Recover it withPOST /v1/apps/{appId}/deploywhen it has areadyversion, or with a newappSourcethroughPATCH /v1/apps/{appId}. Lifecycle is not health.activesays the app is deployed and taking work, not that workers exist to run it: an app whose workers the platform cannot currently place staysactive. Readhealthfor that.Possible values7 values
isFavorite
booleanrequiredWhether the authenticated organization has favorited this app. Favorited apps sort ahead of non-favorited apps. Toggled via
PUT/DELETE/v1/apps/{appId}/favorite.
activeVersionId
string (uuid)nullableUUID v4Current deployed version. Null until the first version is successfully deployed.
createdAt
string (date-time)requireddate-time
updatedAt
string (date-time)requireddate-time
Errors
| Status | When |
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Summary across apps
Aggregate dashboard metrics across all apps owned by the authenticated organization. App and worker tallies are always present. Request and error-rate totals come from the metrics store and are omitted when that hop cannot answer rather than reported as zero. Spend covers a rolling 24 hours and is omitted when the usage cannot be priced.
Response
activeApps
integerrequiredint64Apps currently in the
activestatus.
totalApps
integerrequiredint64All apps in the organization excluding soft-deleted ones.
activeWorkers
integerrequiredint64Non-terminal workers (
statusother thanstopped) across every app and version in the organization.
provisionedGpuCount
integerrequiredint64Sum of
gpuCountacross those same non-terminal workers.
requests24h
integerint64Requests served in the last 24 hours across every app. Omitted when metrics cannot be read. Zero when the organization had no traffic.
errorRate24h
numberdoubleError ratio (4xx + 5xx over requests) for the last 24 hours, in 0–1. Omitted when metrics cannot be read or when
requests24his zero.
spendToday
objectProvisional pay-as-you-go accrual over the last 24 hours, across every app. Excludes finalization rounding and ledger adjustments. A rolling window rather than a calendar day, so the figure does not reset at a boundary the customer did not choose. What has accrued, not a projection: an app that has just started shows what it has run so far. Time a capacity commitment covered is excluded, because it collects nothing. Omitted when the usage cannot be priced, which a zero would misreport as having run for free.
calculatedAt
string (date-time)requireddate-timeWhen these metrics were computed.
Errors
| Status | When |
401 | Missing or invalid credentials |
500 | Unexpected server error |
503 | A required service is temporarily unavailable |