Guaranteed GPUs for a term
Commit to a number of GPUs of a given type for a term, billed monthly or annually at a negotiated rate. Those GPUs are yours for the duration. Anything above the reservation bursts into pay as you go.
Bring your code or container and run it on Runware. We handle GPU scheduling, workers, autoscaling and execution from zero to production capacity.
See workers, GPU allocation, queue depth, errors, usage and cost across every Serverless app.
Runware Serverless apps dashboard listing active apps with workers, GPUs, latency, error rates, and 24-hour request trends.
Serverless is for teams running custom, fine-tuned or multimodal workloads who need more control than a hosted model API without managing GPU infrastructure themselves.
Bring your code for the simplest path to deploy, or use your own container when you need full control over the runtime and dependencies.
Provide your workload code and let Runware package it for you, or deploy your own container when you need full runtime control.
01 · Package
Choose the GPU type and the number of GPUs each worker needs.
02 · Compute
Choose minimum and maximum workers. Scale fully to zero when idle, or keep workers warm for latency-sensitive workloads.
03 · Scale
See worker state, logs, GPU usage, requests and cost from the Serverless console.
04 · Observe
Bring a container or upload your workload code and get a production endpoint that scales with traffic. We handle the GPU infrastructure, scaling, routing and worker lifecycle.
Upload, configure, deploy. Step through the flow and your workload is live on an autoscaling endpoint.
Purpose-built infrastructure for running demanding GPU workloads without building the platform underneath them.
Reduce startup time with memory and process snapshotting, backed by an optimized filesystem designed for fast model and runtime loading.
Respond instantly as demand changes. Scale from zero to large GPU fleets without provisioning or managing the infrastructure yourself.
Get real-time visibility across workers, GPUs, requests and performance with granular metrics, logs and insights built into the platform.
Run everything from image and video generation to multimodal pipelines, LLM inference and custom GPU apps on the infrastructure that best fits the job.
Pay as you go for elasticity. Reserve capacity for certainty. Most production teams combine the two: a reserved baseline with pay-as-you-go burst above it.
Commit to a number of GPUs of a given type for a term, billed monthly or annually at a negotiated rate. Those GPUs are yours for the duration. Anything above the reservation bursts into pay as you go.
Special launch pricing, metered per second. GPU prices fluctuate, so these list rates are subject to change. Talk to us to lock in today's rate for a committed term. Multi-GPU layouts and GB300 are scoped with our engineers before any commitment.
You pay while a worker is allocated to your app. Holding workers warm with a minimum worker count is still pay as you go, not a reservation.
Time spent waiting for a GPU, pulling your image or recovering from a failure on our side is not metered.
GPU usage is metered by the second and attributed to each Serverless app.
Runware serverless usage dashboard: per-second metering across apps with a 30-day cost-over-time breakdown, per-app GPU-seconds and cost, prepaid credit balance and auto top-up settings.
Serverless operates the workload for you. If you need your own Kubernetes stack, SSH access or dedicated infrastructure, use GPU Compute.
Bring a container image or handler code for GPU-accelerated workloads including image, video and diffusion pipelines, fine-tuned and multimodal models, LLM inference and other custom AI applications.
Choose the GPU type and number of GPUs each worker needs, with support for both single and multi-GPU workloads.
Yes. Serverless supports multi-GPU workers for workloads that need more memory or compute than a single GPU can provide.
You can configure the number of GPUs assigned to each worker, and Runware schedules those GPUs together as a single workload unit. Supported layouts depend on the GPU type and available capacity, with common configurations including 2, 4 and 8 GPUs per worker.
Multi-GPU is suited to workloads such as large-model inference, high-memory generation pipelines and applications that already support frameworks such as NCCL, tensor parallelism or pipeline parallelism.
Because all GPUs for a worker must be available together, multi-GPU workloads may take longer to schedule than single-GPU workloads when shared capacity is constrained.
When an app has no work to process, Runware can scale its workers down to zero so you are not paying for idle GPU capacity.
When a new request arrives, Serverless allocates suitable GPU capacity, starts a worker and begins processing the queued workload. You control the minimum and maximum number of workers for each app, so workloads that need consistently warm capacity can keep workers running while more elastic workloads can scale fully to zero.
GPU usage is billed only once capacity has been allocated to your workload.
When an app scales from zero, Runware needs to allocate a GPU and prepare a worker before the first request can run. Startup time depends on factors such as GPU availability, container size, model size and how much data needs to be loaded into GPU memory.
Serverless is designed to minimize this overhead through image caching, workload placement and checkpointing. Checkpointing allows compatible workers to preserve prepared workload state so they can be restored more efficiently rather than repeating the full startup process each time.
For latency-sensitive workloads, you can also configure a minimum number of workers to remain warm, trading some idle GPU cost for more predictable response times.
Pay as you go for variable traffic, development and burst. Workers come from the shared pool and are billed per second while allocated. Reserved when production cannot miss capacity: a fixed number of GPUs for a term, guaranteed. Most production teams reserve a baseline and burst above it on pay as you go.
No. A minimum worker count keeps GPUs allocated from the shared pool and bills at the pay-as-you-go rate for the whole time. Reserved capacity is a commercial commitment that guarantees the GPUs even when the pool is full.
Create a Runware account, then request access from the Serverless overview in your dashboard. We review fit and capacity, enable Serverless on your workspace, and set up an onboarding or a short proof of concept with agreed success criteria.
Deploy custom AI workloads without building the GPU platform underneath them.