Monitoring
Read an app's events and failed requests, and the named metric and log queries behind them.
Introduction
Two shapes of read. Events and failed requests hang off an app and are listed directly. Metrics and logs go through named queries: you ask the catalog what it can answer, then read one query by its id.
Monitoring covers which of these answers which question. This page is the routes.
Reading logs, metrics and an app's request errors draws on a read allowance for your organization. When it runs out those routes answer 429 with a Retry-After, and a client that ignores it keeps being refused.
List app events
Events are the platform's record of what it did: deploy, scaling, audit and error. scaling is usually the one to read first when an app costs more or answers slower than you expected, because it says when workers appeared and disappeared.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
Response
nextCursor
stringnullableCursor for the next page. Null when there are no more items.
data
object[]Array items7 properties each
id
string (uuid)requiredUUID v4
appId
stringrequiredmin: 6max: 30Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches
deleted.
workerId
string (uuid)nullableUUID v4
endpointId
string (uuid)nullableUUID v4
type
stringrequiredPossible values4 values
message
stringrequiredHuman-readable description of what happened. Never blank.
createdAt
string (date-time)date-time
Errors
| Status | When |
400 | The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
|
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
500 | Unexpected server error |
503 | A required service is temporarily unavailable |
List request errors
One page of failed inference requests for this app, newest first. Omit statusClass for both 4xx and 5xx. The cursor is opaque and is only valid with the same window and statusClass it was issued under.
Omitting statusClass returns 4xx and 5xx together, which is usually what you want first: the split between them is the diagnosis. The cursor is only valid with the window and statusClass it was issued under, so changing either means starting again.
Request
appId
stringrequiredmin: 6max: 30Immutable app identifier, unique among the authenticated organization's live apps.
limit
integerint32min: 1max: 100default: 20Maximum number of items to return.
cursor
stringOpaque pagination cursor returned by a previous call, as
nextCursoror, on the operations that offer one,prevCursor.
window
stringdefault: 24hThe range to search. Same closed ladder as the metrics queries. Defaults to the last 24 hours.
Allowed values5 values
statusClass
stringNarrow to one error class. Omit for both 4xx and 5xx.
Allowed values2 values
Response
Errors
| Status | When |
400 | The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
|
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
429 | The organization's read allowance is exhausted. Retry-After says how long to wait. A client that ignores it will keep being refused.
|
500 | Unexpected server error |
502 | An upstream service was unreachable or refused the request |
503 | A required service is temporarily unavailable |
504 | An upstream service did not answer in time |
The query catalog
Every metric and log read names a query the platform already knows how to answer, and the catalog is the source of those ids: discover them from it rather than carrying a list of your own.
Each entry also reports the windows it can answer. A window missing from that list has no stored series behind it, so offering it would render an empty chart.
Some queries are deliberately absent from the catalog because they back a specific surface: the runtime object on an endpoint, and an app's overview. Reading one by id still works, so an absence from the catalog is about what a client should offer rather than what it can ask for.
List queries
The catalog: every named query, its unit and aggregation, the selectors it accepts, the series it returns, and the windows actually backed by stored series.
This is the source of query ids for the Metrics and Logs tabs. Clients discover ids here rather than carrying a list of their own, and render window tabs from windows rather than from the full ladder, so a window whose storage tier has no backing series stays invisible instead of rendering a tab with nothing behind it.
The list sparklines apps_request_volume and endpoints_request_volume are the only list-page queries here. The other list-page runtime queries and the app overview queries are not charts, so this catalog omits them. getMetricSeries serves them by id.
Response
metrics
object[]requiredArray items6 properties each
id
stringrequiredThe value to pass as
queryId.
unit
stringrequiredThe unit of every value in this query's series, e.g.
req/min.
aggregation
stringrequiredHow each bucket was reduced, e.g.
avgorp50. Carried so a tooltip can say what a point means rather than presenting a bucket reduction as an instant reading.
windows
string[]requiredThe windows this query can answer. A window absent here has no stored series behind it, so a client should not offer it.
selectors
string[]requiredThe narrowings this query accepts. Anything else is rejected.
series
string[]requiredThe stable series ids this query returns. Empty when those ids are not known until the request (
apps_request_volumekeys each line on an app id).
logs
object[]requiredArray items6 properties each
id
stringrequiredThe value to pass as
queryId.
unit
stringrequiredThe unit of every value in this query's series, e.g.
req/min.
aggregation
stringrequiredHow each bucket was reduced, e.g.
avgorp50. Carried so a tooltip can say what a point means rather than presenting a bucket reduction as an instant reading.
windows
string[]requiredThe windows this query can answer. A window absent here has no stored series behind it, so a client should not offer it.
selectors
string[]requiredThe narrowings this query accepts. Anything else is rejected.
series
string[]requiredThe stable series ids this query returns. Empty when those ids are not known until the request (
apps_request_volumekeys each line on an app id).
Errors
| Status | When |
401 | Missing or invalid credentials |
500 | Unexpected server error |
502 | An upstream service was unreachable or refused the request |
503 | A required service is temporarily unavailable |
504 | An upstream service did not answer in time |
Read a metric query
Returns one chart's data: a single timestamp axis shared by every series, and one dense value array per series aligned to it.
Values are dense and positionally aligned to t, with an explicit null wherever there was no sample. Each timestamp is the END of its bucket, so window.to is inclusive and equals the last timestamp in t, while window.from is exclusive and is one step_s before the first.
An organization with no metrics yet is answered with the full axis and all-null series rather than an error.
apps_request_volume returns one series per app: 96 quarter-hour request counts over window=24h (step_s 900, unit requests). Repeat appId once per id on the current list page to pad idle apps with all-null series, in request order. A series is named for the live app behind it, so a reused app id reports its own generation's traffic and not the one before it. The same appId pad applies to the list-scoped apps_error_volume and apps_request_duration queries. Those feed listApps / getApp runtime and are omitted from the catalog. Other queries reject appId. Other windows are not available for these queries.
endpoints_request_volume is the endpoints-list counterpart: 96 quarter-hour request counts per endpoint over window=24h (step_s 900, unit requests). It requires the public app id. Repeat endpointId once per id on the current listEndpoints page to pad idle endpoints with all-null series, in request order. Buckets that started before that endpoint row's createdAt are null, so a removed-then-readded path does not inherit the previous row's traffic.
The same app + endpointId pad and createdAt clip apply to the list-scoped endpoints_request_rate, endpoints_latency_p95, endpoints_latency_p99, endpoints_error_volume and endpoints_request_count queries (window=1h only, 5-minute step). Those feed listEndpoints / getEndpoint runtime and are omitted from the catalog. Other queries reject endpointId. Other windows are not available for these queries.
throughput, latency and error_rate accept endpoint (the allocated UUID) together with the public app id, so the Metrics-tab charts can be scoped to one path. Other cataloged queries reject endpoint. endpoint is a matcher, not the endpointId pad.
gpu_utilisation, gpu_power_usage, gpu_power_utilisation, gpu_temperature and gpu_memory_used require the public app id and accept worker (Worker.id, the Kubernetes pod UID), so the Metrics tab can chart an app and a worker page can chart one live worker. Other cataloged queries reject worker. gpu_count stays app-scoped.
Three queries serve one app's overview, all over window=24h at step_s 900 and all requiring the public app id: app_traffic_24h returns requests, client_errors (4xx) and server_errors (5xx) as request counts. app_worker_seconds_24h returns startup, execution and idle as worker-seconds, whose three values in a bucket sum to that bucket's worker time, and app_cold_starts_24h returns cold_starts as a count. They report counts and totals rather than rates or ratios, so a per-minute figure is a bucket value divided by step_s / 60 and a 24h ratio is one summed axis over another, summing first and dividing once, because averaging a per-bucket ratio across the axis does not give the 24h ratio. These queries are absent from listInsightsQueries: they back the overview rather than the Metrics tab. Other windows are not available for them.
One chart per read: a single timestamp axis shared by every series, and one dense value array per series aligned to it, with an explicit null wherever there was no sample. Each timestamp is the end of its bucket, so a point at t[i] covers the step_s seconds before it. An organization with no metrics yet is answered with the full axis and all-null series rather than an error.
Request
queryId
stringrequiredA query id from the catalog. Deliberately not an enum: the downstream registry is the source of ids, and an enum here would be a second list to keep in step with it.
window
stringrequiredThe time window. A closed set rather than a free-form range, because every distinct range defeats the server-side cache alignment that makes a sliding window cheap. Only the windows a query lists in the catalog can be asked of it.
Allowed values5 values
pinnedTo
integerint64Fixes the window's inclusive end, the last timestamp in
t, so that several calls making up one visual share an axis instead of racing the clock between them. Must be aligned to the window's step, no newer than the newest readable edge, and inside retention.
deployment
stringmin: 6max: 30Narrow to one app.
endpoint
string (uuid)UUID v4Narrow to one endpoint within an app. The value is the allocated
endpoints.id(UUID), never the customer-authored path. Accepted bythroughput,latencyanderror_rate, and requires the public app id. A query that does not declare this selector rejects it rather than ignoring it.
statusClass
stringNarrow to one response class.
Allowed values3 values
region
stringNarrow to one region. No series carries a region label yet, so no query currently accepts this and supplying it is rejected rather than ignored.
worker
string (uuid)UUID v4Narrow to one live worker within an app. The value is
Worker.id(the Kubernetes pod UID). Accepted bygpu_utilisation,gpu_power_usage,gpu_power_utilisation,gpu_temperatureandgpu_memory_used, and requires the public app id. A query that does not declare this selector rejects it rather than ignoring it.
appId
string[]Restrict the list-scoped queries (
apps_request_volume,apps_error_volume,apps_request_duration) to these app ids: one series per id, in request order, with all-null series for apps that had no samples. Values are quarter-hourly overwindow=24h. An id with no live app is dropped rather than refused: a deleted app has no traffic to pad. Repeat the parameter once per id on the current list page (at most 100, matchinglistApps). Omit it to receive every app in the organization that had data, unless that set is larger than this query will expand: then the call is a422onappIdand the list page should name the apps it is showing. Other queries reject this parameter.
endpointId
string (uuid)[]Restrict expand-by-
endpoint_idqueries (endpoints_request_volume,endpoints_request_rate,endpoints_latency_p95,endpoints_latency_p99,endpoints_error_volume,endpoints_request_count) to these endpoint ids: one series per id, in request order, with all-null series for endpoints that had no traffic.endpoints_request_volumeis quarter-hourly request counts overwindow=24h. The others are latest-window scalars overwindow=1h. Buckets that started before that endpoint row'screatedAtare null. Repeat the parameter once per id on the current list page (at most 100, matchinglistEndpoints). Omit it to receive every endpoint on the selected app that had data, unless that set is larger than this query will expand: then the call is a422onendpointIdand the list page should name the endpoints it is showing. Other queries reject this parameter. Requires the public app id.
Response
query
stringrequired
unit
stringrequired
step_s
integerrequiredint64Seconds between adjacent timestamps in
t.
aggregation
stringrequired
window
objectrequiredThe range covered, as
(from, to].
t
integer[]requiredThe timestamp axis, shared by every series. Each value is the END of its bucket, in unix seconds: a point at
t[i]is the aggregate over(t[i] - step_s, t[i]], which is what the underlying range functions compute.
series
object[]requiredArray items3 properties each
id
stringrequiredStable identifier for this line. For chart queries it does not change, so a client can map it to a color or legend position. For
apps_request_volumeit is the app id of that line.
label
stringrequiredDisplay text. May change without notice. Do not switch on it.
v
number[]requiredValues, positionally aligned to
tand always the same length.nullis a genuine absence of data, never a zero, and is sent explicitly rather than omitted so a gap draws as a gap instead of a line across an outage.
Errors
| Status | When |
400 | The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
|
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
429 | The organization's read allowance is exhausted. Retry-After says how long to wait. A client that ignores it will keep being refused.
|
500 | Unexpected server error |
502 | An upstream service was unreachable or refused the request |
503 | A required service is temporarily unavailable |
504 | An upstream service did not answer in time |
Read log entries
Returns one page of log entries in the requested sort order, newest first by default, with opaque cursors for the neighboring pages when they exist. nextCursor continues in the sort order and prevCursor goes back against it, so a client can walk a window in either direction from either end. A cursor is only valid for the sort it was issued under. Reusing one under the other ordering returns 400. Query ids and their supported selectors are listed by the insights catalog.
nextCursor continues in the sort order and prevCursor goes back against it, so you can walk a window from either end. A cursor is only valid for the sort it was issued under.
Request
queryId
stringrequiredA query id from the catalog. Deliberately not an enum: the downstream registry is the source of ids, and an enum here would be a second list to keep in step with it.
window
stringrequiredThe time window. A closed set rather than a free-form range, because every distinct range defeats the server-side cache alignment that makes a sliding window cheap. Only the windows a query lists in the catalog can be asked of it.
Allowed values5 values
limit
integerint32min: 1max: 100default: 20Maximum number of items to return.
cursor
stringOpaque pagination cursor returned by a previous call, as
nextCursoror, on the operations that offer one,prevCursor.
sort
stringOrdering of a log page and the direction
nextCursormoves in. OnlygetLogEntriesaccepts it.Allowed values2 values
deployment
stringmin: 6max: 30Narrow to one app.
endpoint
string (uuid)UUID v4Narrow to one endpoint within an app. The value is the allocated
endpoints.id(UUID), never the customer-authored path. Accepted bythroughput,latencyanderror_rate, and requires the public app id. A query that does not declare this selector rejects it rather than ignoring it.
Response
entries
object[]required
nextCursor
stringOpaque. Continues in the
sortorder. Absent on the last page.
prevCursor
stringOpaque. Goes back against the
sortorder to the page before this one. Absent on the first page. Inside a group of entries sharing one timestamp, which has no order of its own, the pages read backwards can split the group differently from the pages read forwards. Each entry is still returned once either way.
Errors
| Status | When |
400 | The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
|
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
429 | The organization's read allowance is exhausted. Retry-After says how long to wait. A client that ignores it will keep being refused.
|
500 | Unexpected server error |
502 | An upstream service was unreachable or refused the request |
503 | A required service is temporarily unavailable |
504 | An upstream service did not answer in time |
Follow logs
Streams new application log entries as Server-Sent Events. The stream sends keepalive comments while quiet and ends with an end event when its connection lifetime expires or the service shuts down. Clients should reconnect after an end event. Use the runtime_tail query id. Queries that sort or aggregate are rejected because they cannot be followed live.
Server-Sent Events, with keepalive comments while the stream is quiet. It ends with an end event when its connection lifetime expires, and the client is expected to reconnect. A live stream has no window, and a query that sorts or aggregates is rejected because it cannot be followed.
Request
queryId
stringrequiredA query id from the catalog. Deliberately not an enum: the downstream registry is the source of ids, and an enum here would be a second list to keep in step with it.
deployment
stringrequiredmin: 6max: 30App whose new log entries are streamed.
Errors
| Status | When |
401 | Missing or invalid credentials |
404 | Resource not found |
422 | The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
|
429 | The organization's read allowance is exhausted. Retry-After says how long to wait. A client that ignores it will keep being refused.
|
500 | Unexpected server error |
502 | An upstream service was unreachable or refused the request |
503 | A required service is temporarily unavailable |
504 | An upstream service did not answer in time |