Models/Collections/Build realtime generation products
Blueprint4 ModelsUpdated May 2026

Build realtime generation products

Sub-second image generation and realtime speech for interactive apps, copilots, and live product experiences.

Realtime stack.Generation and speech at interactive speed

GenerateSub-secondPhotoreal

Z-Image-Turbo

Fast photorealistic image generator with text control

$0.0034/megapixel
Run

Z-Image-Turbo, a distilled model for sub-second image generation. Sharp photoreal output with accurate English and Chinese text inside images.

TXT → IMGIMG → IMG

Tradeoff. Priced around $0.0034 per megapixel ($0.0013 at 512x512, $0.0032 at 1024x1024); 2048x2048 jumps to $0.0141 per image.

Open weights.Hosted alternatives for this stack

Open-weight models covering the same tasks as this stack, running on Runware's own optimized compute and billed on compute time. License terms vary per model — check each model page before self-hosting.

Black Forest Labs

FLUX.2 [klein] 9B KV

by Black Forest Labs

KV-cache accelerated image generation and editing for real-time multi-reference workflows

TXT → IMGIMG → IMG
from $0.00078/img
Black Forest Labs

FLUX.2 [dev]

by Black Forest Labs

FLUX.2 dev for controllable open text to image workflows

TXT → IMGIMG → IMG
from $0.0051/512x512
RunDiffusion

Juggernaut Z

by RunDiffusion

Polished image model with stronger cinematic lighting, cleaner focus, and richer portrait detail

TXT → IMGIMG → IMG
from $0.0117
Black Forest Labs

FLUX.2 [klein] 9B

by Black Forest Labs

Ultra-fast image generation and editing with sub-second latency

TXT → IMGIMG → IMG
from $0.00078/1024x1024

Common questions

Generate — Z-Image-Turbo; Generate — P-Image; Speak — Inworld Realtime TTS-2; Enhance — Llama 3.1 8B Prompt Enhancer. Each row above expands with the model's live pricing, capability chips, and sample outputs.

Z-Image-Turbo and Llama 3.1 8B Prompt Enhancer run on Runware's own optimized compute (the platform's open-weight tier, billed on compute time) — check each model page for weight availability and license terms before self-hosting. The remaining picks are partner-served.

Tradeoffs to consider

Z-Image-Turbo and P-Image both commit to sub-second generation in their own catalog descriptions, and Inworld TTS-2 is built for low-latency realtime interaction. The catalog publishes no measured latency figures, so validate against your own regions and payload sizes before promising an SLA.

Llama 3.1 8B Prompt Enhancer is currently the only model in the Best Prompt Enhance collection and ships no published per-call price. Treat it as optional in the hot path until you have verified its cost and added a fallback (for example, passing the raw prompt through).

Production notes

Every price on this page is the model's published rate from the live Runware catalog, using the cheapest listed configuration unless stated otherwise. Prices vary with resolution, duration, quality tier, or token volume, so check the pricing table on each model page before estimating unit economics.

The Runware catalog does not publish per-model latency figures, so this page does not quote end-to-end timings. Where a model's own description commits to speed (for example sub-second generation or realtime streaming), that claim is repeated here. For anything else, benchmark the exact models in the Playground with your own payload sizes before committing to an SLA.

Models are addressed by versioned AIR identifiers, so a workflow pinned to specific model versions keeps producing the same behaviour as new versions ship. Adopt upgrades deliberately by re-running your evaluation set against the new version before switching production traffic.

About this collection

A blueprint bundles the small set of models you would actually wire together for one production use case, instead of ranking a whole category. Each pick covers one stage of the workflow and links to its model page for full schema, pricing, and examples.

Picks are live models from the Fastest Image Generation (17 models), Fastest Audio Generation (7), and Best Prompt Enhance (1) collections, selected for explicit realtime or sub-second positioning in their own catalog descriptions.

Model guides.Learn how to use the stack

Inworld Realtime TTS-2

Formatting LLM output for speech

How to write LLM system prompts that produce text TTS-2 can synthesize naturally, with normalization, filler words, and emphasis cues handled before the audio call.

Read the guide →
Inworld Realtime TTS-2

Controlling voice delivery with steering tags

How to use natural-language steering tags to control emotion, pacing, volume, and vocal style in TTS-2 speech output.

Read the guide →