Workers
Inspect the workers currently serving an application.
Introduction
A worker is one running copy of your app, and it serves one task at a time, so the number of workers is the whole of your throughput.
The states these routes report are where a cold start becomes visible: a worker waiting for capacity is not the same as one whose pod came up and never became ready. Monitoring covers what each state means and which one to act on.
List workers
Returns a newest-first page of workers observed for the app (including terminal stopped rows until purged). Omitted versionId scopes the page to the app's activeVersionId. An app with no active version therefore answers an empty default page, not that it has no workers, only that none are pinned. Optional state, status and q narrow the page further. A cursor must be replayed under the same filters it was issued with.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
limit
integerint32min: 1max: 100default: 20Maximum number of items to return.
cursor
stringOpaque pagination cursor returned by a previous call, as
nextCursoror, on the operations that offer one,prevCursor.
versionId
stringmin: 1Scope the page to one version. The default is the app's
activeVersionId. When that field is unset the default page is empty: the app has no active version, not that it has no workers. Sendall(any case) to include every version. An empty value is refused. A cursor must be replayed under the same version scope it was issued with.
state
stringdefault: allNarrow the page by worker state. The default,
all, keeps the terminalstoppedrows in the page.livedrops them. Astateoflivewith astatusofstoppedis a contradiction and is refused, because an empty page would read as "this app has never run".Allowed values2 values
status
stringWorker lifecycle status. Also the type of
UsageEvent.eventType, which records a ledger subset of these states (see that field:busynever appears there).unhealthymeans the pod exists but failed to become or stay ready, not an intentional drain or stop.Allowed values9 values
q
stringmin: 1max: 100Case-insensitive literal substring match against
id,podName,nodeNameandversionId. A worker matching any field is returned within the selected version, state and status filters. Omitqto disable search. Cursors must retain the same search term, ignoring case.
Response
nextCursor
stringnullableCursor for the next page. Null when there are no more items.
data
object[]Array items14 properties each
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
versionId
string (uuid)requiredUUID v4
podName
stringrequired
nodeName
stringnullableKubernetes node the pod is scheduled on. Omitted while still pending.
status
stringrequiredWorker lifecycle status. Also the type of
UsageEvent.eventType, which records a ledger subset of these states (see that field:busynever appears there).unhealthymeans the pod exists but failed to become or stay ready, not an intentional drain or stop.Possible values9 values
statusReason
stringnullableKubernetes-facing reason for the current status when unhealthy or otherwise notable (e.g.
ImagePullBackOff,CrashLoopBackOff). Apendingworker the cluster could not place reportsUnschedulable, which means no node could give it the GPUs the app asks for.
statusOccurredAt
string (date-time)requireddate-timeWhen the worker entered its current status. For a
stoppedworker this is when it stopped, so it is also the end of the worker's uptime.
gpuCount
integerrequiredint32min: 0GPUs attached to this worker at observation time. Zero means the snapshot is unknown: the worker is still unscheduled, or a terminal observation could not read a GPU pair. A known snapshot is
>= 1and travels withgpuType. An empty pair is incomplete, not a CPU worker. CPU workloads are not supported.
gpuType
stringnullablemin: 1max: 64GPU catalog code snapshotted at observation time. Omitted until a GPU snapshot exists (the worker is still unscheduled, or the observation could not read a type). Not a CPU-worker marker. CPU workloads are not supported.
gpuAvailability
stringnullableHow the catalog currently provisions this worker's
gpuType. Omitted when there is nogpuType, or when the catalog no longer holds the code. UnlikegpuTypethis is read now rather than snapshotted, so it tells you how that GPU is supplied today, not how it was supplied when the worker started.Possible values3 values
lastSeenAt
string (date-time)nullabledate-timeWorker-reported liveness. Omitted until the worker reports. Stays absent until a heartbeat path exists (phase 2).
createdAt
string (date-time)requireddate-timePod creation time (
metadata.creationTimestamp), not insert time. It is also the start of the worker's uptime.
updatedAt
string (date-time)date-time
Errors
| Status | When |
400 | The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
|
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
Get a worker
Returns one worker by id within the app. The id is the Kubernetes pod UID recorded by the reconciler. A worker that belongs to another app (or tenant) is not found.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
workerId
string (uuid)requiredUUID v4Worker identifier, the Kubernetes pod UID recorded by the reconciler. This is the stable key to address a worker with.
podNameis the value to show a person.
Response
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
versionId
string (uuid)requiredUUID v4
podName
stringrequired
nodeName
stringnullableKubernetes node the pod is scheduled on. Omitted while still pending.
status
stringrequiredWorker lifecycle status. Also the type of
UsageEvent.eventType, which records a ledger subset of these states (see that field:busynever appears there).unhealthymeans the pod exists but failed to become or stay ready, not an intentional drain or stop.Possible values9 values
statusReason
stringnullableKubernetes-facing reason for the current status when unhealthy or otherwise notable (e.g.
ImagePullBackOff,CrashLoopBackOff). Apendingworker the cluster could not place reportsUnschedulable, which means no node could give it the GPUs the app asks for.
statusOccurredAt
string (date-time)requireddate-timeWhen the worker entered its current status. For a
stoppedworker this is when it stopped, so it is also the end of the worker's uptime.
gpuCount
integerrequiredint32min: 0GPUs attached to this worker at observation time. Zero means the snapshot is unknown: the worker is still unscheduled, or a terminal observation could not read a GPU pair. A known snapshot is
>= 1and travels withgpuType. An empty pair is incomplete, not a CPU worker. CPU workloads are not supported.
gpuType
stringnullablemin: 1max: 64GPU catalog code snapshotted at observation time. Omitted until a GPU snapshot exists (the worker is still unscheduled, or the observation could not read a type). Not a CPU-worker marker. CPU workloads are not supported.
gpuAvailability
stringnullableHow the catalog currently provisions this worker's
gpuType. Omitted when there is nogpuType, or when the catalog no longer holds the code. UnlikegpuTypethis is read now rather than snapshotted, so it tells you how that GPU is supplied today, not how it was supplied when the worker started.Possible values3 values
lastSeenAt
string (date-time)nullabledate-timeWorker-reported liveness. Omitted until the worker reports. Stays absent until a heartbeat path exists (phase 2).
createdAt
string (date-time)requireddate-timePod creation time (
metadata.creationTimestamp), not insert time. It is also the start of the worker's uptime.
updatedAt
string (date-time)date-time
Errors
| Status | When |
400 | The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
|
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |