Models/Collections/Best Text-to-Speech
Capability · Audio19 ModelsUpdated Jun 2026

Best Text-to-Speech

Models that generate natural-sounding speech with good pacing, pronunciation, and distinct voice character. Useful for narration, virtual assistants, and production-ready voice output.

The picks.Natural voice synthesis

TXT → AUD

Fish Audio S2.1 Pro

Flagship multilingual text-to-speech with natural language voice control and realtime streaming

TXT → AUDNative audio80+ langsStreamingMulti-speaker

Common questions

Yes — ACE-Step v1.5 XL SFT, ACE-Step v1.5 XL Base, ACE-Step v1.5 XL Turbo, Qwen3-TTS 1.7B VoiceDesign and 5 more run on Runware's own optimized compute (the platform's open-weight tier, billed on compute time). Check each model page for license terms before self-hosting.

Fish Audio S2.1 Pro, released June 2026 per the live catalog. Membership updates automatically as the catalog publishes new models to this collection.

Production notes

Every price on this page is the model's published rate from the live Runware catalog, using the cheapest listed configuration unless stated otherwise. Prices vary with resolution, duration, quality tier, or token volume, so check the pricing table on each model page before estimating unit economics.

The Runware catalog does not publish per-model latency figures, so this page does not quote end-to-end timings. Where a model's own description commits to speed (for example sub-second generation or realtime streaming), that claim is repeated here. For anything else, benchmark the exact models in the Playground with your own payload sizes before committing to an SLA.

Models are addressed by versioned AIR identifiers, so a workflow pinned to specific model versions keeps producing the same behaviour as new versions ship. Adopt upgrades deliberately by re-running your evaluation set against the new version before switching production traffic.

About this collection

Models that generate natural-sounding speech with good pacing, pronunciation, and distinct voice character. Useful for narration, virtual assistants, and production-ready voice output.

Membership comes directly from the Runware catalog: the 19 models on this page are the live catalog's own membership for the "Best Text-to-Speech" collection. Names, descriptions, pricing, capability chips, samples, and guides are all read live from the catalog — nothing here is hand-curated.

Model guides.Learn how to use the stack

Fish Audio S2.1 Pro

Emotion and expression control

How to control vocal delivery in Fish Audio S2-Pro with bracket tags. The tag system steers emotion, expression, paralanguage, and phoneme-level pronunciation in one inline syntax.

Read the guide →
Fish Audio S2.1 Pro

Multi-speaker dialogue

How to generate two-speaker dialogue audio in a single request to Fish Audio S2-Pro using inline speaker tags. One call, two voices, full per-speaker emotion control.

Read the guide →
Inworld Realtime TTS-2

Formatting LLM output for speech

How to write LLM system prompts that produce text TTS-2 can synthesize naturally, with normalization, filler words, and emphasis cues handled before the audio call.

Read the guide →
Inworld Realtime TTS-2

Controlling voice delivery with steering tags

How to use natural-language steering tags to control emotion, pacing, volume, and vocal style in TTS-2 speech output.

Read the guide →