Monitoring

Read an app's events and failed requests, and the named metric and log queries behind them.

Introduction

Two shapes of read. Events and failed requests hang off an app and are listed directly. Metrics and logs go through named queries: you ask the catalog what it can answer, then read one query by its id.

Monitoring covers which of these answers which question. This page is the routes.

Reading logs, metrics and an app's request errors draws on a read allowance for your organization. When it runs out those routes answer 429 with a Retry-After, and a client that ignores it keeps being refused.

List app events

GETapi.serverless.runware.ai/v1/apps/{appId}/events

Events are the platform's record of what it did: deploy, scaling, audit and error. scaling is usually the one to read first when an app costs more or answers slower than you expected, because it says when workers appeared and disappeared.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Query

limit

integerint32min: 1max: 100default: 20

Maximum number of items to return.

cursor

string

Opaque pagination cursor returned by a previous call, as nextCursor or, on the operations that offer one, prevCursor.

type

string
Allowed values4 values

Response

nextCursor

stringnullable

Cursor for the next page. Null when there are no more items.

data

object[]
Array items7 properties each

id

string (uuid)requiredUUID v4

appId

stringrequiredmin: 6max: 30

Immutable app identifier. Unique among the authenticated organization's live apps: it cannot be changed after creation, and it becomes available again once the app it named reaches deleted.

workerId

string (uuid)nullableUUID v4

endpointId

string (uuid)nullableUUID v4

type

stringrequired
Possible values4 values

message

stringrequired

Human-readable description of what happened. Never blank.

createdAt

string (date-time)date-time

Errors

StatusWhen
400The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
500Unexpected server error
503A required service is temporarily unavailable

List request errors

GETapi.serverless.runware.ai/v1/apps/{appId}/errors

One page of failed inference requests for this app, newest first. Omit statusClass for both 4xx and 5xx. The cursor is opaque and is only valid with the same window and statusClass it was issued under.

Omitting statusClass returns 4xx and 5xx together, which is usually what you want first: the split between them is the diagnosis. The cursor is only valid with the window and statusClass it was issued under, so changing either means starting again.

Request

Path

appId

stringrequiredmin: 6max: 30

Immutable app identifier, unique among the authenticated organization's live apps.

Query

limit

integerint32min: 1max: 100default: 20

Maximum number of items to return.

cursor

string

Opaque pagination cursor returned by a previous call, as nextCursor or, on the operations that offer one, prevCursor.

window

stringdefault: 24h

The range to search. Same closed ladder as the metrics queries. Defaults to the last 24 hours.

Allowed values5 values

Narrow to one error class. Omit for both 4xx and 5xx.

Allowed values2 values

Response

nextCursor

stringnullable

Cursor for the next page. Null when there are no more items.

data

object[]required
Array items4 properties each

time

integerrequiredint64

level

string

body

stringrequired

fields

object

Errors

StatusWhen
400The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
429The organization's read allowance is exhausted. Retry-After says how long to wait. A client that ignores it will keep being refused.
500Unexpected server error
502An upstream service was unreachable or refused the request
503A required service is temporarily unavailable
504An upstream service did not answer in time

The query catalog

Every metric and log read names a query the platform already knows how to answer, and the catalog is the source of those ids: discover them from it rather than carrying a list of your own.

Each entry also reports the windows it can answer. A window missing from that list has no stored series behind it, so offering it would render an empty chart.

Some queries are deliberately absent from the catalog because they back a specific surface: the runtime object on an endpoint, and an app's overview. Reading one by id still works, so an absence from the catalog is about what a client should offer rather than what it can ask for.

List queries

GETapi.serverless.runware.ai/v1/metrics/queries

The catalog: every named query, its unit and aggregation, the selectors it accepts, the series it returns, and the windows actually backed by stored series. This is the source of query ids for the Metrics and Logs tabs. Clients discover ids here rather than carrying a list of their own, and render window tabs from windows rather than from the full ladder, so a window whose storage tier has no backing series stays invisible instead of rendering a tab with nothing behind it. The list sparklines apps_request_volume and endpoints_request_volume are the only list-page queries here. The other list-page runtime queries and the app overview queries are not charts, so this catalog omits them. getMetricSeries serves them by id.

Response

metrics

object[]required
Array items6 properties each

id

stringrequired

The value to pass as queryId.

unit

stringrequired

The unit of every value in this query's series, e.g. req/min.

aggregation

stringrequired

How each bucket was reduced, e.g. avg or p50. Carried so a tooltip can say what a point means rather than presenting a bucket reduction as an instant reading.

windows

string[]required

The windows this query can answer. A window absent here has no stored series behind it, so a client should not offer it.

selectors

string[]required

The narrowings this query accepts. Anything else is rejected.

series

string[]required

The stable series ids this query returns. Empty when those ids are not known until the request (apps_request_volume keys each line on an app id).

logs

object[]required
Array items6 properties each

id

stringrequired

The value to pass as queryId.

unit

stringrequired

The unit of every value in this query's series, e.g. req/min.

aggregation

stringrequired

How each bucket was reduced, e.g. avg or p50. Carried so a tooltip can say what a point means rather than presenting a bucket reduction as an instant reading.

windows

string[]required

The windows this query can answer. A window absent here has no stored series behind it, so a client should not offer it.

selectors

string[]required

The narrowings this query accepts. Anything else is rejected.

series

string[]required

The stable series ids this query returns. Empty when those ids are not known until the request (apps_request_volume keys each line on an app id).

Errors

StatusWhen
401Missing or invalid credentials
500Unexpected server error
502An upstream service was unreachable or refused the request
503A required service is temporarily unavailable
504An upstream service did not answer in time

Read a metric query

GETapi.serverless.runware.ai/v1/metrics/queries/{queryId}/series

Returns one chart's data: a single timestamp axis shared by every series, and one dense value array per series aligned to it. Values are dense and positionally aligned to t, with an explicit null wherever there was no sample. Each timestamp is the END of its bucket, so window.to is inclusive and equals the last timestamp in t, while window.from is exclusive and is one step_s before the first. An organization with no metrics yet is answered with the full axis and all-null series rather than an error. apps_request_volume returns one series per app: 96 quarter-hour request counts over window=24h (step_s 900, unit requests). Repeat appId once per id on the current list page to pad idle apps with all-null series, in request order. A series is named for the live app behind it, so a reused app id reports its own generation's traffic and not the one before it. The same appId pad applies to the list-scoped apps_error_volume and apps_request_duration queries. Those feed listApps / getApp runtime and are omitted from the catalog. Other queries reject appId. Other windows are not available for these queries. endpoints_request_volume is the endpoints-list counterpart: 96 quarter-hour request counts per endpoint over window=24h (step_s 900, unit requests). It requires the public app id. Repeat endpointId once per id on the current listEndpoints page to pad idle endpoints with all-null series, in request order. Buckets that started before that endpoint row's createdAt are null, so a removed-then-readded path does not inherit the previous row's traffic. The same app + endpointId pad and createdAt clip apply to the list-scoped endpoints_request_rate, endpoints_latency_p95, endpoints_latency_p99, endpoints_error_volume and endpoints_request_count queries (window=1h only, 5-minute step). Those feed listEndpoints / getEndpoint runtime and are omitted from the catalog. Other queries reject endpointId. Other windows are not available for these queries. throughput, latency and error_rate accept endpoint (the allocated UUID) together with the public app id, so the Metrics-tab charts can be scoped to one path. Other cataloged queries reject endpoint. endpoint is a matcher, not the endpointId pad. gpu_utilisation, gpu_power_usage, gpu_power_utilisation, gpu_temperature and gpu_memory_used require the public app id and accept worker (Worker.id, the Kubernetes pod UID), so the Metrics tab can chart an app and a worker page can chart one live worker. Other cataloged queries reject worker. gpu_count stays app-scoped. Three queries serve one app's overview, all over window=24h at step_s 900 and all requiring the public app id: app_traffic_24h returns requests, client_errors (4xx) and server_errors (5xx) as request counts. app_worker_seconds_24h returns startup, execution and idle as worker-seconds, whose three values in a bucket sum to that bucket's worker time, and app_cold_starts_24h returns cold_starts as a count. They report counts and totals rather than rates or ratios, so a per-minute figure is a bucket value divided by step_s / 60 and a 24h ratio is one summed axis over another, summing first and dividing once, because averaging a per-bucket ratio across the axis does not give the 24h ratio. These queries are absent from listInsightsQueries: they back the overview rather than the Metrics tab. Other windows are not available for them.

One chart per read: a single timestamp axis shared by every series, and one dense value array per series aligned to it, with an explicit null wherever there was no sample. Each timestamp is the end of its bucket, so a point at t[i] covers the step_s seconds before it. An organization with no metrics yet is answered with the full axis and all-null series rather than an error.

Request

Path

queryId

stringrequired

A query id from the catalog. Deliberately not an enum: the downstream registry is the source of ids, and an enum here would be a second list to keep in step with it.

Query

window

stringrequired

The time window. A closed set rather than a free-form range, because every distinct range defeats the server-side cache alignment that makes a sliding window cheap. Only the windows a query lists in the catalog can be asked of it.

Allowed values5 values

pinnedTo

integerint64

Fixes the window's inclusive end, the last timestamp in t, so that several calls making up one visual share an axis instead of racing the clock between them. Must be aligned to the window's step, no newer than the newest readable edge, and inside retention.

deployment

stringmin: 6max: 30

Narrow to one app.

endpoint

string (uuid)UUID v4

Narrow to one endpoint within an app. The value is the allocated endpoints.id (UUID), never the customer-authored path. Accepted by throughput, latency and error_rate, and requires the public app id. A query that does not declare this selector rejects it rather than ignoring it.

Narrow to one response class.

Allowed values3 values

region

string

Narrow to one region. No series carries a region label yet, so no query currently accepts this and supplying it is rejected rather than ignored.

worker

string (uuid)UUID v4

Narrow to one live worker within an app. The value is Worker.id (the Kubernetes pod UID). Accepted by gpu_utilisation, gpu_power_usage, gpu_power_utilisation, gpu_temperature and gpu_memory_used, and requires the public app id. A query that does not declare this selector rejects it rather than ignoring it.

appId

string[]

Restrict the list-scoped queries (apps_request_volume, apps_error_volume, apps_request_duration) to these app ids: one series per id, in request order, with all-null series for apps that had no samples. Values are quarter-hourly over window=24h. An id with no live app is dropped rather than refused: a deleted app has no traffic to pad. Repeat the parameter once per id on the current list page (at most 100, matching listApps). Omit it to receive every app in the organization that had data, unless that set is larger than this query will expand: then the call is a 422 on appId and the list page should name the apps it is showing. Other queries reject this parameter.

endpointId

string (uuid)[]

Restrict expand-by-endpoint_id queries (endpoints_request_volume, endpoints_request_rate, endpoints_latency_p95, endpoints_latency_p99, endpoints_error_volume, endpoints_request_count) to these endpoint ids: one series per id, in request order, with all-null series for endpoints that had no traffic. endpoints_request_volume is quarter-hourly request counts over window=24h. The others are latest-window scalars over window=1h. Buckets that started before that endpoint row's createdAt are null. Repeat the parameter once per id on the current list page (at most 100, matching listEndpoints). Omit it to receive every endpoint on the selected app that had data, unless that set is larger than this query will expand: then the call is a 422 on endpointId and the list page should name the endpoints it is showing. Other queries reject this parameter. Requires the public app id.

Response

query

stringrequired

unit

stringrequired

step_s

integerrequiredint64

Seconds between adjacent timestamps in t.

aggregation

stringrequired

window

objectrequired

The range covered, as (from, to].

Properties2 properties

from

integerrequiredint64

EXCLUSIVE: one step_s before the first timestamp in t. Nothing is sampled at from itself.

to

integerrequiredint64

INCLUSIVE, and equal to the last timestamp in t.

t

integer[]required

The timestamp axis, shared by every series. Each value is the END of its bucket, in unix seconds: a point at t[i] is the aggregate over (t[i] - step_s, t[i]], which is what the underlying range functions compute.

series

object[]required
Array items3 properties each

id

stringrequired

Stable identifier for this line. For chart queries it does not change, so a client can map it to a color or legend position. For apps_request_volume it is the app id of that line.

label

stringrequired

Display text. May change without notice. Do not switch on it.

v

number[]required

Values, positionally aligned to t and always the same length. null is a genuine absence of data, never a zero, and is sent explicitly rather than omitted so a gap draws as a gap instead of a line across an outage.

Errors

StatusWhen
400The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
429The organization's read allowance is exhausted. Retry-After says how long to wait. A client that ignores it will keep being refused.
500Unexpected server error
502An upstream service was unreachable or refused the request
503A required service is temporarily unavailable
504An upstream service did not answer in time

Read log entries

GETapi.serverless.runware.ai/v1/logs/queries/{queryId}/entries

Returns one page of log entries in the requested sort order, newest first by default, with opaque cursors for the neighboring pages when they exist. nextCursor continues in the sort order and prevCursor goes back against it, so a client can walk a window in either direction from either end. A cursor is only valid for the sort it was issued under. Reusing one under the other ordering returns 400. Query ids and their supported selectors are listed by the insights catalog.

nextCursor continues in the sort order and prevCursor goes back against it, so you can walk a window from either end. A cursor is only valid for the sort it was issued under.

Request

Path

queryId

stringrequired

A query id from the catalog. Deliberately not an enum: the downstream registry is the source of ids, and an enum here would be a second list to keep in step with it.

Query

window

stringrequired

The time window. A closed set rather than a free-form range, because every distinct range defeats the server-side cache alignment that makes a sliding window cheap. Only the windows a query lists in the catalog can be asked of it.

Allowed values5 values

limit

integerint32min: 1max: 100default: 20

Maximum number of items to return.

cursor

string

Opaque pagination cursor returned by a previous call, as nextCursor or, on the operations that offer one, prevCursor.

sort

string

Ordering of a log page and the direction nextCursor moves in. Only getLogEntries accepts it.

Allowed values2 values

deployment

stringmin: 6max: 30

Narrow to one app.

endpoint

string (uuid)UUID v4

Narrow to one endpoint within an app. The value is the allocated endpoints.id (UUID), never the customer-authored path. Accepted by throughput, latency and error_rate, and requires the public app id. A query that does not declare this selector rejects it rather than ignoring it.

Response

entries

object[]required
Array items4 properties each

time

integerrequiredint64

level

string

body

stringrequired

fields

object

Opaque. Continues in the sort order. Absent on the last page.

Opaque. Goes back against the sort order to the page before this one. Absent on the first page. Inside a group of entries sharing one timestamp, which has no order of its own, the pages read backwards can split the group differently from the pages read forwards. Each entry is still returned once either way.

Errors

StatusWhen
400The request was malformed and could not be parsed (e.g. invalid JSON). A well-formed request that fails validation returns 422 instead.
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
429The organization's read allowance is exhausted. Retry-After says how long to wait. A client that ignores it will keep being refused.
500Unexpected server error
502An upstream service was unreachable or refused the request
503A required service is temporarily unavailable
504An upstream service did not answer in time

Follow logs

GETapi.serverless.runware.ai/v1/logs/queries/{queryId}/tail

Streams new application log entries as Server-Sent Events. The stream sends keepalive comments while quiet and ends with an end event when its connection lifetime expires or the service shuts down. Clients should reconnect after an end event. Use the runtime_tail query id. Queries that sort or aggregate are rejected because they cannot be followed live.

Server-Sent Events, with keepalive comments while the stream is quiet. It ends with an end event when its connection lifetime expires, and the client is expected to reconnect. A live stream has no window, and a query that sorts or aggregates is rejected because it cannot be followed.

Request

Path

queryId

stringrequired

A query id from the catalog. Deliberately not an enum: the downstream registry is the source of ids, and an enum here would be a second list to keep in step with it.

Query

deployment

stringrequiredmin: 6max: 30

App whose new log entries are streamed.

Errors

StatusWhen
401Missing or invalid credentials
404Resource not found
422The request was well-formed but semantically invalid (e.g. a missing or out-of-range field). A request that could not be parsed returns 400. The errors array carries one entry per offending field.
429The organization's read allowance is exhausted. Retry-After says how long to wait. A client that ignores it will keep being refused.
500Unexpected server error
502An upstream service was unreachable or refused the request
503A required service is temporarily unavailable
504An upstream service did not answer in time