Monitoring
Read the workers serving an app, the traffic on each endpoint, the events behind a scaling decision, and the requests that failed.
Introduction
Five surfaces tell you what an app is doing, and most questions need just one of them. Health is the verdict, in one field. Workers are the present tense. Endpoint traffic is the recent past, per path. Events are the control plane's record of what the platform did. Errors are the failed requests themselves.
runware serverless apps workers my-app
runware serverless apps events my-appWhether the app can serve at all
Reading an app carries a health object, and it is a different question from status. Status is what you asked for. Health is what the cluster managed. An app can read active and serve nothing.
| State | What it means |
healthy | Can serve |
degraded | Can serve, with less capacity than it asked for |
unavailable | Has nothing able to serve |
The reason says why, and only two of the six are the platform being short of GPUs. capacity_below_floor means the app keeps a warm floor and has fewer workers than that floor. demand_unserved means the app scales to zero and has work queued with nothing running it. Those two are the ones that refuse an invocation with capacity-unavailable.
The other four are not capacity. workload_present, autoscaler_unhealthy, workload_terminating and workload_missing describe a workload that is fine, unmanaged, going away or gone, and a request refused for one of those gets the plain service-unavailable.
health is absent until the platform has observed the app, and absence is not a claim that it is healthy. Read it as a report that lags the cluster by one observation rather than as a prediction of whether your next call will land.
An initializing app that already has a version to route to accepts invocations while reporting unavailable, which is the ordinary case during a first deploy. Only an active app is refused on health.
Where a worker is in its life
A worker moves through a fixed sequence, and the state it is stuck in tells you what is slow.
| State | What is happening |
pending | Waiting for capacity. Nothing of yours is running yet. A statusReason of Unschedulable means no node could give it the GPUs the app asks for |
pulling | Fetching the image. Large images pay here |
loading | Your load is running. Weights that are not on a volume pay here |
ready | Idle and able to take work |
busy | Serving |
unhealthy | The pod exists but never became ready, or stopped being ready |
draining | Finishing in-flight work before it goes |
stopping, stopped | Shutting down, or gone |
That sequence is the cold start, broken into its parts. A worker sitting in pulling says your image is too big. A worker sitting in loading says your weights are being fetched, which is what a volume fixes. The state tells you which.
unhealthy is worth reading precisely: it means the pod came up and did not become ready, not that something drained or stopped it deliberately. For a container app that usually points at the readiness probe.
Traffic, per endpoint
Listing an app's endpoints returns more than paths. Each row may carry a runtime object holding the latest five-minute window of that endpoint's traffic, so a single call says which path is busy and which one is slow.
GET /v1/apps/{appId}/endpoints| Field | What it reports |
status | A badge derived from the error ratio: healthy at zero, degraded below 5%, unhealthy at 5% and above |
requestsPerMinute | The five-minute request rate |
p95RequestDuration | 95th percentile request duration, in seconds, across all requests |
p99RequestDuration | 99th percentile, same basis |
The object is omitted rather than zeroed when the endpoint had no requests in the window, or when the metrics store cannot be read. A zero that is present is a reading: it means no traffic was measured, not that nothing is known.
status is computed, not stored. There is no status column on an endpoint and nothing sets one, so it reflects the last five minutes and nothing longer. An endpoint that failed badly an hour ago and has been quiet since reads as absent, not as unhealthy.
Durations cover all requests regardless of outcome, because the duration histogram carries no status class. A p99 that climbs while the error rate climbs with it may be measuring how long failures take.
apps endpoints lists the paths. The runtime numbers come from the API read above.
Events
Events are the record of what the platform did and why, and they are filterable by kind.
| Type | Covers |
deploy | Rollouts, activations, rollbacks |
scaling | Workers added and removed, and what triggered it |
audit | Configuration and attachment changes |
error | Failures the platform recorded against the app |
runware serverless apps events my-app --type scalingScaling events are the ones worth reading first when an app costs more or responds slower than you expected. They say when workers appeared and disappeared, which is the difference between a scaling policy that is wrong and traffic that is spikier than you thought.
A deploy event also names the endpoint paths a deploy adds and removes, because renaming a handler method moves a public URL. The deploy goes ahead and the event records the change, so it is where you find out that yesterday's callers are now getting a 404.
What the GPU is doing
Four charts read the GPU itself: power draw, power utilization, temperature and memory used.
They take a worker selector, so you can follow one pod, and without it they answer with the mean across the app's workers. Metric charts are read through the query catalog rather than by writing your own: see the Monitoring API.
Memory used against the card's capacity is the one to check when a model loads on one GPU type and fails on a smaller one.
Failed requests
Events record what the platform did. Errors record what your callers got. One page of failed inference requests, newest first, is available per app.
GET /v1/apps/{appId}/errors?window=24h&statusClass=5xxwindow takes 1h, 6h, 24h, 7d or 30d and defaults to the last day. Omitting statusClass returns both 4xx and 5xx together, which is usually what you want first: the split between the two is the diagnosis. A wall of 4xx is callers sending the wrong shape, and a wall of 5xx is your handler.
Each entry carries a timestamp, a level, the message body and a set of string fields.
The cursor is opaque and only valid with the window and statusClass it was issued under. Changing either mid-pagination means starting again.
What to look at, in order
When an app is not answering, the questions have an order and each one rules out a layer.
Is it deployed? An app still building has no workers by definition. Check its status before anything else.
Are there workers, and what state? No workers on an app scaled to zero is correct and not a fault. Workers stuck in pending mean capacity, and unhealthy means your image came up and did not become ready.
Is one endpoint the problem, or all of them? The runtime object on each endpoint answers that in one call, and it is the fastest way to tell a broken path from a broken app.
Did a deploy change something? deploy and audit events cover both the code and the configuration, and a configuration-only change still records a version you can compare.
Is the request even reaching an endpoint? A 404 carrying endpointPath never reached a worker at all. See Errors.
Logs and handler errors
Read what your workers print with apps logs, and --follow to watch it live. See the CLI reference.
An app that fails inside your handler surfaces its reason on the task's error, and that string includes the message your own exception raised. Make it say something useful. The failed-request listing above tells you that a call failed and when, not what your code was doing at the time.