Pricing
You pay for GPU time, not for requests. What that means for a scaled-to-zero app, a busy one, and the gap between them.
Introduction
You pay for GPU time, and nothing else meters. A worker holding a GPU costs money from the moment it holds it until it lets go. Requests, tasks, endpoints and everything else about your app are free.
That single sentence explains most of what surprises people, so it is worth taking literally before reading further.
The unit
Cost is the time a worker held a GPU, multiplied by how many GPUs it held, at the rate for that GPU type.
runware serverless gpusThat lists the available types with their per-second price. Rates are in USD.
The rate is recorded on each usage event as it happens, so changing an app's GPU type or repricing a type leaves earlier charges as they are. What you paid is what the rate was at the time.
Time is rounded up to the whole second, once per GPU, so a very short worker life still costs a minimum slice.
Serverless credit
Serverless draws on its own prepaid credit balance, separate from the balance your inference requests use. GPU time is charged against it as it settles. Starting capacity needs credit, see Errors.
A balance that goes negative, which a reversed top-up or a chargeback can do, retires the workers that credit was paying for. Workers covered by reserved capacity keep running, because that capacity is already paid for. Clearing the debt lifts it with nothing for you to do.
Reserved capacity
Pay as you go draws from a shared pool, so it is subject to availability. Reserved capacity is the other way to pay for the same GPUs: you commit to a number of them for a term and they stay yours when the pool is full, with anything above the reservation falling back to pay as you go.
It is arranged rather than self-served. Talk to sales to put one in place, and read what a reservation covered through the usage API.
Where the time goes
A worker's life splits into three, and all three are GPU time you pay for.
| Part | What is happening |
| Startup | Running your load, from the moment your container starts. The GPU is held and nothing is being served. Waiting for capacity and fetching the image are free |
| Execution | Serving requests. The part you actually wanted |
| Idle | Ready and waiting for the next request |
The platform reports the split, so this is measurable. Startup and idle are where an unexamined app spends money it did not mean to.
A worker that is not serving can still be charged
The rule is whether the worker holds a GPU, not whether it is doing anything useful with it.
| State | Charged |
pending, pulling | No. There is no GPU held yet |
loading, ready | Yes. The GPU is held whether or not a request is in flight |
draining, stopping | Yes, until the worker is gone |
unhealthy | Only when your container is crash-looping. An image that fails to pull, a lost node and a node that went not-ready are not charged |
stopped | No |
A crash-looping container bills for as long as it keeps restarting, because it holds its GPU on every attempt. An app stuck that way is the most expensive thing on this page, and nothing stops it for you. Watch for unhealthy in Monitoring.
Scale to zero
minWorkers at 0 means an app with no traffic has no workers and costs no GPU time at all. The first request after a quiet period starts one, and pays for that start.
Holding minWorkers above zero buys away the cold start, and you pay for that capacity whether or not anything arrives. That trade is the main one this product asks of you, and either side of it can be the right call.
idleTtlSecs is the quiet lever. It is how long a worker survives with nothing to do, and every second of it is billed. A long TTL makes the second request after a gap fast and pays for the gap. A short one saves the gap and pays a cold start more often.
What costs nothing
| Thing | Why it is free |
| A rejected invocation | A 404, 422 or 503 is refused before a task exists, so no worker runs |
| Resubmitting a task id | Returns the task it already names rather than starting a second run |
| A configuration change | Records a version and carries the image forward, so it costs no build |
| A rollback | Re-applies an image that version already pinned. Nothing is rebuilt |
| A stopped app | Has released its workers. It holds its name, its versions and its volumes |
| A refused deploy | A 402 is decided before anything starts, and holds no credit |
Making it cheaper
Put your weights on a volume. This is the largest single lever most apps have. The sandbox filesystem is part of the checkpointed state, so weights left there are fetched again on every cold start, and every one of those downloads is GPU time at your GPU's rate.
Watch where a cold start actually goes. A worker stuck in pulling means the image is too big, which costs latency alone. Stuck in loading means the weights are being fetched, and that time is billed. Monitoring breaks the sequence into its parts, so you can see which one you are paying for.
Check whether the idle is doing anything for you. If your traffic arrives in bursts an hour apart, a 15 minute idleTtlSecs is paying for 15 minutes of nothing, several times a day.
Right-size the GPU rather than the count. Every endpoint on an app shares one worker pool, so an app whose cheapest endpoint forces an expensive GPU is paying that rate for all of them. Two endpoints with genuinely different hardware needs are two apps.
Reading what you spent
Usage is an append-only ledger of worker state transitions, one row per transition, each carrying the GPU count and the rate that applied.
GET /v1/usage?appId=my-appBilling is drawn on when a transition happened, not when it was recorded, so ingest lag never moves an event across a boundary. A worker's transitions stay listed for at least 120 days after its last one.
busy never appears in this ledger. It describes queue occupancy, and this ledger records changes in the worker's life. A worker costs the same whether it is serving or waiting. That is the whole pricing model restated: the meter follows the pod, not the request.
For totals, GET /v1/usage/summary returns GPU time and spend over a window of up to 31 days, grouped by the dimensions you ask for.
GET /v1/usage/summary?appId=my-apppaygSpend is provisional. It excludes the final rounding and ledger adjustments, so it can differ slightly from what is settled.
From the terminal, runware serverless usage prints the same summary, and runware serverless apps usage <appId> narrows it to one app. See the CLI reference.