Modal alternative

Serverless GPUs for up to 79% less than Modal.

Run your Python or your container on GPUs in modular data centres we design and build ourselves. Per-second billing, and reserved capacity at a published rate from $0.63 per GPU-hour.

From $0.63Per GPU-hour, reserved
Up to 79%Lower than Modal
1,000 GPUsScale to, then back to zero

Every GPU both platforms sell, list price to list price. On some GPUs the Modal serverless rate and Runware Serverless are close; RTX PRO 6000 is 34% lower and L40S is 18% lower with no commitment. Reserved capacity, which Modal publishes no rate for, is 12–79% lower across the board on a 1 to 24 month term.

Runware Serverless vs Modal
USD per GPU-hour. Both platforms bill by the second; neither GPU rate includes CPU or memory.
GPUModalRunware ServerlessSavingRunware reservedSaving
RTX PRO 600096 GB
$3.03$1.99−34%
as low as$0.63
−79%
B300280 GB
$7.10$7.00Within 1%
as low as$3.43
−52%
L40S48 GB
$1.95$1.60−18%
as low as$0.99
−49%
H10080 GB
$3.95$3.95Same price
as low as$2.45
−38%
H200141 GB
$4.54$4.50Within 1%
as low as$3.00
−34%
B200180 GB
$6.25$6.00−4%
as low as$5.49
−12%

GPU list prices per GPU-hour from modal.com/pricing and runware.ai/serverless, 6 Oct 2026. Modal publishes per-second rates; hourly figures are those × 3,600. Runware reserved rates require a 1–24 month term. Billed per second. GPU prices fluctuate, so list rates are subject to change.

Worked example

One RTX PRO 6000 for a month, around the clock

What 730 hours on a single RTX PRO 6000 costs per month, at the list rates in the table above. Reserve a baseline and burst above it on Serverless to land between the two Runware figures.

$2,212Modal
$1,453Runware Serverless
$460Runware reserved
Why teams switch

The same shape of product: Python in, a scaled endpoint out. The difference is underneath it, from a published reserved rate to the data centre the GPUs sit in.

Reserved capacity at a published price
Commit to a number of GPUs for 1 to 24 months at rates as low as $0.63 per GPU-hour, listed on our Serverless page. Modal's pricing page lists no reserved or committed rate.
Dedicated data-centre GPUs we run ourselves
Serverless runs on Sonic Pods, the modular data centres Runware designs and operates. You choose the exact GPU type, and that is the GPU your worker gets.
Per-second billing, only while allocated
You pay for the time a worker holds a GPU. Rejected invocations, resubmitted task ids, configuration changes, rollbacks and stopped apps cost nothing.
Scale to 1,000 GPUs, then back to zero
Set minimum and maximum workers per app. Scale fully to zero when idle and back up as traffic arrives, without provisioning anything yourself.
Built for Python-first teams
Decorate a class, deploy a file. A CLI (runware auth login, runware serverless deploy), Python and TypeScript SDKs, and a load method that runs once per worker before any request arrives.
Inference, training and fine-tuning
Image, video and diffusion pipelines, LLM inference, fine-tuning and custom GPU apps, on single or multi-GPU workers of 2, 4 or 8 GPUs.
Four regions
Deploy in us-east, us-west, eu-west or eu-central, and keep workloads close to your users and your data.
One account, one bill
The same workspace also gets Runware's Models API: ready-made endpoints for open image, video, audio and language models next to Serverless for your own.
Migration

One thing moves: Modal puts the GPU and the image in your source file, and Runware keeps them on the deploy, so a configuration change never needs a rebuild. Everything else is a rename.

  1. Step 01

    Bring your class, or your container

    Most of a Modal app translates line for line. @app.cls becomes @serve, @modal.enter() becomes a load method, and each web endpoint becomes @endpoint. If your deployment is a FastAPI or Flask app, bring the container instead: your image runs the server and a container.yaml declares the routes.

  2. Step 02

    Move the hardware out of the decorator

    GPU type, min and max workers, pip requirements and volumes become flags on the deploy, and stay changeable on the live app from the CLI or dashboard without a rebuild. Set --max-workers explicitly: it defaults to 1 here.

  3. Step 03

    Point traffic at the new endpoint

    Callers send a POST with a taskId and your fields inside payload; the same taskId twice returns the same task, so retries are safe. Watch workers, logs and cost in the console, then wind the Modal app down.

  4. Migrate to RunwareServerless quickstartModal migration guide
model.pyRunware Serverless
from runware_serverless import endpoint, serve

@serve
class ImageTools:
    def load(self) -> None:
        self.pipe = load_my_pipeline()

    @endpoint
    def generate(self, prompt: str) -> dict:
        return {"image": self.pipe(prompt).encode()}
Deploy
runware auth login
runware serverless deploy model.py --id image-tools --gpu-type h100 \
  --requirement torch --requirement diffusers
A Runware Sonic Pod: a white 20 ft container with liquid-cooling units on top, in a field beside a solar array.
Powered by Sonic Pods

Why the price is lower: we build the data centre too.

Runware Serverless runs on Sonic Pods, modular AI data centres we design, build and operate ourselves. Up to 1,200 directly liquid-cooled GPUs in a 20 ft container, with no oversized building, no overbuilt redundancy and no years of sunk cost in front of them. Every component inside exists to serve inference, so you pay for GPU time rather than someone else's overhead.

Cheaper buildout vs traditional
100×
Power efficiency, liquid cooled
95%
To deploy a pod, not years
Days
No trainingOn your data
SSO + SAMLEnterprise auth
SOC 2Certified
ISO 27001Certified
GDPRCompliant
24/7Engineering support

Teams that switched and scaled

Lower bill, no GPU platform to run. Here's what moving to Runware looked like for teams shipping at production scale.

HeyGen product interface preview
HeyGen logo

HeyGen

How a leading AI video platform unlocked new product capabilities by replacing expensive managed providers with a single, cost-effective inference API.

OpenArt case study preview
OpenArt logo

OpenArt

How OpenArt achieved sub-second image generation at scale with Runware's Sonic Engine.

Higgsfield case study preview
Higgsfield logo

Higgsfield AI

How a fast-growing video generation startup deployed a major model release in hours.

NightCafe case study preview
NightCafe logo

NightCafe

How NightCafe scaled their AI art community without hiring a single infrastructure engineer.

Are running with us
FAQ

Yes. Runware Serverless is the same shape of product as Modal serverless: bring Python or a container, and get a GPU endpoint that scales from zero to production capacity and back, billed per second. It runs on data-centre GPUs Runware designs and operates itself, and it sits in the same account as Runware's model inference API.

List price to list price, on the 6 GPUs both platforms sell. With no commitment, RTX PRO 6000 is 34% lower and L40S is 18% lower on Runware Serverless; compare the other GPUs in the table above. Runware reserved capacity, on a 1 to 24 month term, is 12–79% lower than Modal's list price, from $0.63 per GPU-hour; Modal's pricing page lists no reserved rate. Modal prices are from modal.com/pricing on 6 Oct 2026; Runware prices are our published Serverless rates. Neither side's GPU rate includes CPU or memory, which both bill separately.

Serverless draws workers from the shared pool and bills by the second while a worker holds a GPU, with no commitment. Holding workers warm with a minimum worker count is still the Serverless rate. Reserved commits you to a number of GPUs of a given type for a 1 to 24 month term at a lower rate, and those GPUs stay yours even when the pool is full. Anything above the reservation bursts into Serverless. Most production teams reserve a baseline and burst above it.

Both. Serverless is for inference, training and fine-tuning: image, video and diffusion pipelines, LLM inference, fine-tuned and multimodal models and custom GPU apps. Workers can be single-GPU or multi-GPU, with 2, 4 and 8 GPUs per worker, for jobs that need more memory or compute than one card.

Strip the GPU, image and scaling arguments out of @app.cls and hold them for the deploy command. Replace @app.cls with @serve, @modal.enter() with a load method, and each web endpoint decorator with @endpoint; pip_install arguments become --requirement flags. Callers then send a POST with a taskId and the fields inside payload. The Modal migration guide has the full mapping and a porting checklist, and the quickstart takes you from an empty directory to a running endpoint. Our engineers help teams migrate, too.

Reserved capacity

Running at scale? Talk to us about reserved capacity.

Reserved GPUs start as low as $0.63 per GPU-hour on a 1 to 24 month term, and anything above the reservation bursts into Serverless.