Fastest Audio Generation
Models prioritised for speed when generating audio, suitable for rapid iteration and high-throughput production workflows. Useful when latency matters more than maximum output fidelity.
The picks.Real-time synthesis
ACE-Step v1.5 XL Turbo
Fast 4B music generation model with 8-step inference for higher-quality rapid iteration
Qwen3-TTS 1.7B Base
High-quality multilingual text-to-speech with voice cloning and ultra-low latency
Dia2 2B
Streaming dialogue TTS with voice cloning, non-verbal cues, and multi-speaker support
Inworld Realtime TTS-2
Conversational text-to-speech with realtime voice direction and audio-aware delivery
Gemini 3.1 Flash TTS
Expressive text-to-speech with audio tags, multi-speaker dialogue, and 70+ languages
xAI Text-to-Speech
Expressive text-to-speech with over two dozen voices, speech tags, and multilingual support
Fish Audio S2.1 Pro
Flagship multilingual text-to-speech with natural language voice control and realtime streaming
Common questions
Yes — ACE-Step v1.5 XL Turbo, ACE-Step v1.5 Turbo, Qwen3-TTS 1.7B Base, Dia2 2B run on Runware's own optimized compute (the platform's open-weight tier, billed on compute time). Check each model page for license terms before self-hosting.
Fish Audio S2.1 Pro, released June 2026 per the live catalog. Membership updates automatically as the catalog publishes new models to this collection.
Production notes
Every price on this page is the model's published rate from the live Runware catalog, using the cheapest listed configuration unless stated otherwise. Prices vary with resolution, duration, quality tier, or token volume, so check the pricing table on each model page before estimating unit economics.
The Runware catalog does not publish per-model latency figures, so this page does not quote end-to-end timings. Where a model's own description commits to speed (for example sub-second generation or realtime streaming), that claim is repeated here. For anything else, benchmark the exact models in the Playground with your own payload sizes before committing to an SLA.
Models are addressed by versioned AIR identifiers, so a workflow pinned to specific model versions keeps producing the same behaviour as new versions ship. Adopt upgrades deliberately by re-running your evaluation set against the new version before switching production traffic.
About this collection
Models prioritised for speed when generating audio, suitable for rapid iteration and high-throughput production workflows. Useful when latency matters more than maximum output fidelity.
Membership comes directly from the Runware catalog: the 10 models on this page are the live catalog's own membership for the "Fastest Audio Generation" collection. Names, descriptions, pricing, capability chips, samples, and guides are all read live from the catalog — nothing here is hand-curated.
Model guides.Learn how to use the stack
Formatting LLM output for speech
How to write LLM system prompts that produce text TTS-2 can synthesize naturally, with normalization, filler words, and emphasis cues handled before the audio call.
Read the guide →Controlling voice delivery with steering tags
How to use natural-language steering tags to control emotion, pacing, volume, and vocal style in TTS-2 speech output.
Read the guide →Emotion and expression control
How to control vocal delivery in Fish Audio S2-Pro with bracket tags. The tag system steers emotion, expression, paralanguage, and phoneme-level pronunciation in one inline syntax.
Read the guide →Multi-speaker dialogue
How to generate two-speaker dialogue audio in a single request to Fish Audio S2-Pro using inline speaker tags. One call, two voices, full per-speaker emotion control.
Read the guide →