

HeyGen
How a leading AI video platform unlocked new product capabilities by replacing expensive managed providers with a single, cost-effective inference API.
Run your Python or your container on GPUs in modular data centres we design and build ourselves. Per-second billing, and reserved capacity at a published rate from $0.63 per GPU-hour.
Every GPU both platforms sell, list price to list price. On some GPUs the Modal serverless rate and Runware Serverless are close; RTX PRO 6000 is 34% lower and L40S is 18% lower with no commitment. Reserved capacity, which Modal publishes no rate for, is 12–79% lower across the board on a 1 to 24 month term.
GPU list prices per GPU-hour from modal.com/pricing and runware.ai/serverless, 6 Oct 2026. Modal publishes per-second rates; hourly figures are those × 3,600. Runware reserved rates require a 1–24 month term. Billed per second. GPU prices fluctuate, so list rates are subject to change.
What 730 hours on a single RTX PRO 6000 costs per month, at the list rates in the table above. Reserve a baseline and burst above it on Serverless to land between the two Runware figures.
The same shape of product: Python in, a scaled endpoint out. The difference is underneath it, from a published reserved rate to the data centre the GPUs sit in.
One thing moves: Modal puts the GPU and the image in your source file, and Runware keeps them on the deploy, so a configuration change never needs a rebuild. Everything else is a rename.
Most of a Modal app translates line for line. @app.cls becomes @serve, @modal.enter() becomes a load method, and each web endpoint becomes @endpoint. If your deployment is a FastAPI or Flask app, bring the container instead: your image runs the server and a container.yaml declares the routes.
GPU type, min and max workers, pip requirements and volumes become flags on the deploy, and stay changeable on the live app from the CLI or dashboard without a rebuild. Set --max-workers explicitly: it defaults to 1 here.
Callers send a POST with a taskId and your fields inside payload; the same taskId twice returns the same task, so retries are safe. Watch workers, logs and cost in the console, then wind the Modal app down.
from runware_serverless import endpoint, serve
@serve
class ImageTools:
def load(self) -> None:
self.pipe = load_my_pipeline()
@endpoint
def generate(self, prompt: str) -> dict:
return {"image": self.pipe(prompt).encode()}runware auth login
runware serverless deploy model.py --id image-tools --gpu-type h100 \
--requirement torch --requirement diffusers
Runware Serverless runs on Sonic Pods, modular AI data centres we design, build and operate ourselves. Up to 1,200 directly liquid-cooled GPUs in a 20 ft container, with no oversized building, no overbuilt redundancy and no years of sunk cost in front of them. Every component inside exists to serve inference, so you pay for GPU time rather than someone else's overhead.
Lower bill, no GPU platform to run. Here's what moving to Runware looked like for teams shipping at production scale.


How a leading AI video platform unlocked new product capabilities by replacing expensive managed providers with a single, cost-effective inference API.


How OpenArt achieved sub-second image generation at scale with Runware's Sonic Engine.

How a fast-growing video generation startup deployed a major model release in hours.


How NightCafe scaled their AI art community without hiring a single infrastructure engineer.
Yes. Runware Serverless is the same shape of product as Modal serverless: bring Python or a container, and get a GPU endpoint that scales from zero to production capacity and back, billed per second. It runs on data-centre GPUs Runware designs and operates itself, and it sits in the same account as Runware's model inference API.
List price to list price, on the 6 GPUs both platforms sell. With no commitment, RTX PRO 6000 is 34% lower and L40S is 18% lower on Runware Serverless; compare the other GPUs in the table above. Runware reserved capacity, on a 1 to 24 month term, is 12–79% lower than Modal's list price, from $0.63 per GPU-hour; Modal's pricing page lists no reserved rate. Modal prices are from modal.com/pricing on 6 Oct 2026; Runware prices are our published Serverless rates. Neither side's GPU rate includes CPU or memory, which both bill separately.
Serverless draws workers from the shared pool and bills by the second while a worker holds a GPU, with no commitment. Holding workers warm with a minimum worker count is still the Serverless rate. Reserved commits you to a number of GPUs of a given type for a 1 to 24 month term at a lower rate, and those GPUs stay yours even when the pool is full. Anything above the reservation bursts into Serverless. Most production teams reserve a baseline and burst above it.
Both. Serverless is for inference, training and fine-tuning: image, video and diffusion pipelines, LLM inference, fine-tuned and multimodal models and custom GPU apps. Workers can be single-GPU or multi-GPU, with 2, 4 and 8 GPUs per worker, for jobs that need more memory or compute than one card.
Strip the GPU, image and scaling arguments out of @app.cls and hold them for the deploy command. Replace @app.cls with @serve, @modal.enter() with a load method, and each web endpoint decorator with @endpoint; pip_install arguments become --requirement flags. Callers then send a POST with a taskId and the fields inside payload. The Modal migration guide has the full mapping and a porting checklist, and the quickstart takes you from an empty directory to a running endpoint. Our engineers help teams migrate, too.
Reserved GPUs start as low as $0.63 per GPU-hour on a 1 to 24 month term, and anything above the reservation bursts into Serverless.