Best Audio
Frontier voice and music models selected for naturalness, latency, and production readiness: TTS, realtime agents, full-song generation, and open-weight audio.

Flagship multilingual text-to-speech with natural language voice control and realtime streaming

Expressive text-to-speech with audio tags, multi-speaker dialogue, and 70+ languages

Low-latency expressive text-to-speech optimized for real-time apps

Conversational text-to-speech with realtime voice direction and audio-aware delivery

High-quality text-to-speech with expressive, natural voice synthesis

Promptable full-song generation with vocals, lyrics, BPM and key control

4B music generation model with higher audio quality and full editing task support
The picks.Voice, music, and streaming-first synthesis
Production TTS, music, and low-latency voice agents.
Fish Audio S2.1 Pro
Flagship multilingual text-to-speech with natural language voice control and realtime streaming
Flagship multilingual TTS with natural-language voice control and realtime streaming. The default pick for narration, audiobooks, podcasts, and voiceover where stable prosody matters.
Tradeoff. Pricing is measured by input text size (millions of UTF-8 bytes), so budget by script length. For cost-sensitive batch TTS, the lower-priced picks below cover high-volume workloads.
Gemini 3.1 Flash TTS
Expressive text-to-speech with audio tags, multi-speaker dialogue, and 70+ languages
Inworld TTS-1.5 Mini
Low-latency expressive text-to-speech optimized for real-time apps
Inworld Realtime TTS-2
Conversational text-to-speech with realtime voice direction and audio-aware delivery
MiniMax Speech 2.8
High-quality text-to-speech with expressive, natural voice synthesis
MiniMax Music 2.6
Promptable full-song generation with vocals, lyrics, BPM and key control
ACE-Step v1.5 XL Base
4B music generation model with higher audio quality and full editing task support
Common questions
Fish Audio S2.1 Pro is this collection's best-overall pick. ACE-Step v1.5 XL Base is the strongest open-weight option on the list.
Inworld TTS-1.5 Mini is this collection's fastest pick at roughly Low-latency time-to-first-audio. Inworld Realtime TTS-2 is built specifically for stateful multi-turn voice agents. Reserve Fish Audio S2.1 Pro for output that ships to end users.
ACE-Step v1.5 XL Base is this collection's open-weights pick. On Runware it runs on Runware's own optimized compute, billed on compute time. Check the model page for the current license terms before shipping or self-hosting.
Yes. MiniMax Music 2.6 generates full vocal productions — vocals, hooks, arrangement — from a prompt, from $0.15 per 3 min song generation. ACE-Step v1.5 XL Base is the strongest open-weight option for text-to-music workflows.
Tradeoffs to consider
Flagship TTS models (Fish Audio S2.1 Pro, Gemini Flash TTS) produce more nuanced, characterful output at the cost of latency headroom. Realtime-tier models (Inworld TTS 1.5 Mini, Inworld TTS-2) trade some expressive ceiling for low-latency delivery. For live interaction, the latency win is not a compromise: it is the product. For output that ships (audiobooks, narration, ads), expressiveness wins.
TTS and music generation are separate workloads with separate cost structures. TTS is priced per character or per input token. Music generation is priced per second or per song. Avoid comparing them on cost-per-request alone: a 60-second music track and a 60-second narration clip are very different in value and generation complexity.
ACE-Step is the only open-weight pick in this collection and covers music generation only. For open-weight TTS at scale, Alibaba Qwen3 TTS is worth evaluating. Open music generation (ACE-Step) lands within striking distance of frontier MiniMax quality on most prompts: a strong result for a self-hostable pipeline.
All models here are cleared for commercial use through Runware, but underlying license terms vary. Open-weight models carry their own terms on redistribution and training on outputs. Check the model page before any external commercial launch.
Production notes
TTS models bill by input text size, with the unit varying by provider: MiniMax Speech lists per 1,000 characters, Gemini Flash TTS lists per million tokens, and Fish Audio measures millions of UTF-8 bytes of input text. Music models bill per song or per second of output. Each pick row shows the model's own catalog pricing, so plan your spend against your actual script lengths and generation count.
Full pricing reference →For voice agents and interactive apps, streaming time-to-first-audio is the metric that matters, not total generation time. Fish Audio S2.1 Pro ships realtime streaming; Inworld TTS-2 is designed for turn-level streaming with tone carry-forward across turns. If you start with a batch endpoint and switch to streaming later, expect prompt changes and integration work. Wire streaming in your prototype early.
Streaming guide →Every TTS pick in this collection sits behind a Runware-managed endpoint with SLA, but upstream availability varies across providers. Fish Audio S2.1 Pro and MiniMax Speech 2.8 cover the same expressive voice tier. For music generation, MiniMax Music and ACE-Step cover similar creative territory. One-line routing configuration in the Runware API.
Error handling and fallback routing →Pin the model version in your API calls (fishaudio:s2.1@pro, not just a family name). TTS models iterate fast: voice characteristics, language support, and pricing all shift between versions. When a new release lands you want to eval your scripts against it and opt in deliberately.
API versioning reference →Gemini Flash TTS supports inline audio tags in script text and multi-speaker dialogue; xAI TTS ships speech tags. Inworld TTS-2 and Fish Audio S2.1 Pro support free-form natural language voice direction. These control delivery, emotion, and pacing. Invest time in the control layer early.
Voice direction guide →About this collection
The audio generation landscape has moved faster than any other modality in the last eighteen months. This collection answers a single question we hear from product teams building audio features on Runware: which audio models would you ship with today? It covers TTS (expressive, fast, multi-speaker, non-English), music generation (vocal, instrumental, open-weight), and voice agents. Biased toward models with strong naturalness, predictable latency, transparent pricing, and a license you can ship behind.
Every TTS model in the catalog runs a fixed eval set: ten scripts across narration, scripted dialogue, expressive emotional delivery, and multilingual content spanning at least four language families. Scores weight naturalness and prosody most heavily, then intelligibility and emotional accuracy. Latency is measured on realtime-eligible models under actual streaming conditions. Music picks are evaluated on prompt adherence, structural coherence across genres, and production quality.
Every pick links to its model page on Runware with the full API reference, supported parameters, language lists, and pricing examples. The Runware Learning Center covers cross-model audio topics: text normalization for TTS, streaming integration, voice direction and audio tags, LLM-to-speech pipeline design, and music prompt structure.
Browse the Learning Center →New audio models run the eval set the week they land on Runware. When a new entrant clears the bar against the current champion across naturalness, latency, language coverage, and licensing, it replaces the previous pick. Models that fall out of the curated picks remain in the broader audio catalog and appear in the All view.
Every pick in this collection runs on the Runware inference layer: a unified API with predictable latency, transparent pricing, and a single billing surface across TTS and music providers. Routing through Runware means consistent error handling, SLA-backed uptime, version pinning, and the ability to swap between picks without rewriting your integration.
About Runware →Related collections.More ways to build
Browse all collections→Best text to speech
Narration, voice agents, and expressive TTS at every tier.
Open →MUSICBest music generation
Full vocal songs, instrumentals, and music editing.
Open →SPEEDFastest audio generation
Speed-tier picks for realtime interactive audio.
Open →SOTAState of the art models
Frontier production picks across every modality.
Open →