Troubleshooting
A build that produced nothing, a worker that never became ready, a variable that never arrived. What to check, in order.
Introduction
Most problems here are one of five things, and each one shows up in a different place. Knowing which place to look is most of the work.
The build never produced a version
A build moves through queued, building, and then ready or failed. A build that finished as failed recorded no version, so there is nothing to deploy and the app stays where it was.
runware serverless apps builds my-appFor a code app the usual causes are an import missing from your requirements, or an entry file that sits outside the source directory. Import errors surface at build time, because the platform imports your module to read its endpoints out of it.
For a container app, an invalid container.yaml is rejected before a build starts at all: 400 if it cannot be parsed, 422 if it parses and breaks a rule.
A build that shows superseded was not broken. A newer deploy overtook it and the platform canceled it.
The app is active but nothing answers
Check whether there is a worker before assuming there is a fault.
runware serverless apps workers my-appAn app with minWorkers at 0 and no traffic has no workers, and that is correct. The first request starts one, and pays for the start.
If there are workers, their state says what is wrong. pending means the platform is waiting for capacity, and unhealthy means a pod came up and never became ready.
A worker never becomes ready
For a container app this is nearly always the readiness probe. A worker takes traffic only once that route answers.
The mirror image causes a different failure: a server that reports ready before its model is loaded accepts a request it cannot serve, and the caller sees a timeout rather than a cold start.
def do_GET(self):
if self.path == "/ready":
self.send_response(200 if model_loaded else 503)
self.end_headers()Returning 503 until loading finishes is the whole mechanism. See Bringing a container.
The request is rejected
Read the status before the message.
A 404 carrying endpointPath means the app is fine and the path falls outside the set it declares. The check happens before anything is queued, so nothing ran and nothing was charged. A 404 without that member means the app itself is unknown.
A 422 carries one entry per offending field with a JSON pointer. Extra inputs are not permitted means the field belongs to a different endpoint, because endpoint schemas are closed.
A 409 means the app exists but its state cannot take work. A stopped app needs resuming.
A 402 comes from a create, a source update, a deploy, a resume or a scaling change rather than an invocation. Your organization's credit balance cannot back the capacity it would start. shortfall is the amount to add, and the same request succeeds once the balance covers it.
A 503 is the only one worth retrying blindly. No task is created, so a retry is not a duplicate. Full detail in Errors.
The app stops scaling up
effectiveMaxWorkers on the app is the ceiling the last deploy applied. It is lower than configuration.maxWorkers when your credit backed fewer workers at that deploy, and a top-up raises it without a redeploy. It can also differ after a rollback to an older version, or a change made while the app was stopped.
It worked locally and not deployed
Two causes, one per path.
For a container app, the envelope. Runware wraps your payload as {taskId, payload} and hands your server the payload member alone. A local curl sends the inner object directly, so a server that reads the whole body works on your machine and breaks once deployed.
For a code app, optional arguments. An omitted optional field arrives as an explicit null, not as a missing key, so your Python default never fires.
@endpoint
def generate(self, prompt: str, steps: int = 4) -> dict:
steps = 4 if steps is None else stepsThe environment variable did not take
Setting a variable records a new version with the same image and rolls the app, so a running worker picks the value up. Two things stop that.
A write while a rollout is still in flight returns 409 and is discarded. Setting several variables one at a time is that many sequential rollouts with 409s in between, so change them together in one update of the app.
A write that leaves the stored value unchanged records nothing and rolls nothing, which is the other way a variable can look set and change no behavior. See Environment variables.
It is slower than expected
The worker states break a cold start into its parts, so you can see which one is costing you.
A worker sitting in pulling means the image is large. A worker sitting in loading means the weights are being fetched, which is what a volume is for: the sandbox filesystem is part of the checkpointed state, so an unmounted download is copied into every checkpoint and fetched again on every cold start.
If the delay is before either, scaling events say when workers were requested and whether capacity was there.
When the app is up and only some calls are slow, read the endpoints list instead: each row carries its own p95 and p99 for the last five minutes, which separates one slow path from a slow app in a single call.
Reproducing a failure
Worker output is in apps logs, and --follow streams it while you reproduce the problem. The per-app failed-request listing tells you that a call failed and when, and the string on the task's error is the only thing your handler says about it, so make your exceptions say something useful.