Serverless GPUs
for any AI workload

Bring your code or container and run it on Runware. We handle GPU scheduling, workers, autoscaling and execution from zero to production capacity.

From $1.60per GPU-hourPay as you go
NOW LIVE

Everything running. One place.

See workers, GPU allocation, queue depth, errors, usage and cost across every Serverless app.

Runware Serverless apps dashboard listing active apps with workers, GPUs, latency, error rates, and 24-hour request trends.

From workload to running endpoint.

Serverless is for teams running custom, fine-tuned or multimodal workloads who need more control than a hosted model API without managing GPU infrastructure themselves.

Bring your code for the simplest path to deploy, or use your own container when you need full control over the runtime and dependencies.

Serverless containers: deploy your code or a Docker image

Provide your workload code and let Runware package it for you, or deploy your own container when you need full runtime control.

01 · Package

One deploy.
We run the rest.

Bring a container or upload your workload code and get a production endpoint that scales with traffic. We handle the GPU infrastructure, scaling, routing and worker lifecycle.

Upload, configure, deploy. Step through the flow and your workload is live on an autoscaling endpoint.

Upload a container image or your handler code, then configure the runtime environment your workload needs.

Built to move fast.
Built for production.

Purpose-built infrastructure for running demanding GPU workloads without building the platform underneath them.

  • Lightning-fast cold starts

    Reduce startup time with memory and process snapshotting, backed by an optimized filesystem designed for fast model and runtime loading.

    • Memory & process snapshotting
    • Optimized filesystem
    01

Serverless GPU pricing: on-demand or reserved

Pay as you go for elasticity. Reserve capacity for certainty. Most production teams combine the two: a reserved baseline with pay-as-you-go burst above it.

Pay as you go

Per second, from the shared pool

Workers scale up as traffic arrives and back to zero when idle. GPU usage is metered by the second while a worker is allocated to your app. Capacity comes from a shared pool and is subject to availability.
No commitment · from $1.60 per GPU-hour
Reserved capacity

Guaranteed GPUs for a term

Commit to a number of GPUs of a given type for a term, billed monthly or annually at a negotiated rate. Those GPUs are yours for the duration. Anything above the reservation bursts into pay as you go.

1 to 24 month terms · priced per agreement
GPUMemoryBest forPAYG priceReserved
L40s48 GB GDDR6Lighter image and video workloads that fit in 48 GB$1.60/GPU-hour$0.000444 /GPU-secondas low as$0.99/GPU-hour
H10080 GB HBM3LLM inference that fits in 80 GB$3.95/GPU-hour$0.001097 /GPU-secondas low as$2.45/GPU-hour
H200141 GB HBM3eLarger models and long context that no longer fit on H100$4.50/GPU-hour$0.001250 /GPU-secondas low as$3.00/GPU-hour
B200180 GB HBM3eLarge-model Blackwell inference$6.00/GPU-hour$0.001667 /GPU-secondas low as$5.49/GPU-hour
B300280 GB HBM3eFrontier-scale and reasoning-class inference$7.00/GPU-hour$0.001944 /GPU-secondas low as$3.43/GPU-hour
GB300 NVL72Rack-scale Blackwell UltraRack-scale deployments, scoped with our engineersContactScoped with our engineers
CPUPhysical core (2 vCPU equivalent)Pre/post-processing, orchestration and lightweight CPU workloads$0.05/core-hour$0.000013 /core-secondPer agreement

Special launch pricing, metered per second. GPU prices fluctuate, so these list rates are subject to change. Talk to us to lock in today's rate for a committed term. Multi-GPU layouts and GB300 are scoped with our engineers before any commitment.

Billed
  • worker running
  • worker held warm
  • reserved GPUs, used or idle

You pay while a worker is allocated to your app. Holding workers warm with a minimum worker count is still pay as you go, not a reservation.

Not billed
  • scheduling wait
  • image pulls
  • platform failures

Time spent waiting for a GPU, pulling your image or recovering from a failure on our side is not metered.

Know exactly what each workload costs.

GPU usage is metered by the second and attributed to each Serverless app.

Runware serverless usage dashboard: per-second metering across apps with a 30-day cost-over-time breakdown, per-app GPU-seconds and cost, prepaid credit balance and auto top-up settings.

FAQs

Bring a container image or handler code for GPU-accelerated workloads including image, video and diffusion pipelines, fine-tuned and multimodal models, LLM inference and other custom AI applications.

Choose the GPU type and number of GPUs each worker needs, with support for both single and multi-GPU workloads.

Yes. Serverless supports multi-GPU workers for workloads that need more memory or compute than a single GPU can provide.

You can configure the number of GPUs assigned to each worker, and Runware schedules those GPUs together as a single workload unit. Supported layouts depend on the GPU type and available capacity, with common configurations including 2, 4 and 8 GPUs per worker.

Multi-GPU is suited to workloads such as large-model inference, high-memory generation pipelines and applications that already support frameworks such as NCCL, tensor parallelism or pipeline parallelism.

Because all GPUs for a worker must be available together, multi-GPU workloads may take longer to schedule than single-GPU workloads when shared capacity is constrained.

When an app has no work to process, Runware can scale its workers down to zero so you are not paying for idle GPU capacity.

When a new request arrives, Serverless allocates suitable GPU capacity, starts a worker and begins processing the queued workload. You control the minimum and maximum number of workers for each app, so workloads that need consistently warm capacity can keep workers running while more elastic workloads can scale fully to zero.

GPU usage is billed only once capacity has been allocated to your workload.

When an app scales from zero, Runware needs to allocate a GPU and prepare a worker before the first request can run. Startup time depends on factors such as GPU availability, container size, model size and how much data needs to be loaded into GPU memory.

Serverless is designed to minimize this overhead through image caching, workload placement and checkpointing. Checkpointing allows compatible workers to preserve prepared workload state so they can be restored more efficiently rather than repeating the full startup process each time.

For latency-sensitive workloads, you can also configure a minimum number of workers to remain warm, trading some idle GPU cost for more predictable response times.

Pay as you go for variable traffic, development and burst. Workers come from the shared pool and are billed per second while allocated. Reserved when production cannot miss capacity: a fixed number of GPUs for a term, guaranteed. Most production teams reserve a baseline and burst above it on pay as you go.

No. A minimum worker count keeps GPUs allocated from the shared pool and bills at the pay-as-you-go rate for the whole time. Reserved capacity is a commercial commitment that guarantees the GPUs even when the pool is full.

Create a Runware account, then request access from the Serverless overview in your dashboard. We review fit and capacity, enable Serverless on your workspace, and set up an onboarding or a short proof of concept with agreed success criteria.

Serverless · Now live

Bring the workload. We run it.

Deploy custom AI workloads without building the GPU platform underneath them.