Models/Collections/Best Audio
Curated7 ModelsUpdated Jun 2026

Best Audio

Frontier voice and music models selected for naturalness, latency, and production readiness: TTS, realtime agents, full-song generation, and open-weight audio.

The picks.Voice, music, and streaming-first synthesis

Production TTS, music, and low-latency voice agents.

Best overallNarrationAudiobooks

Fish Audio S2.1 Pro

Flagship multilingual text-to-speech with natural language voice control and realtime streaming

Flagship multilingual TTS with natural-language voice control and realtime streaming. The default pick for narration, audiobooks, podcasts, and voiceover where stable prosody matters.

TXT → AUDNative audioMultilingual langsStreamingExpressive

Tradeoff. Pricing is measured by input text size (millions of UTF-8 bytes), so budget by script length. For cost-sensitive batch TTS, the lower-priced picks below cover high-volume workloads.

Common questions

Fish Audio S2.1 Pro is this collection's best-overall pick. ACE-Step v1.5 XL Base is the strongest open-weight option on the list.

Inworld TTS-1.5 Mini is this collection's fastest pick at roughly Low-latency time-to-first-audio. Inworld Realtime TTS-2 is built specifically for stateful multi-turn voice agents. Reserve Fish Audio S2.1 Pro for output that ships to end users.

ACE-Step v1.5 XL Base is this collection's open-weights pick. On Runware it runs on Runware's own optimized compute, billed on compute time. Check the model page for the current license terms before shipping or self-hosting.

Yes. MiniMax Music 2.6 generates full vocal productions — vocals, hooks, arrangement — from a prompt, from $0.15 per 3 min song generation. ACE-Step v1.5 XL Base is the strongest open-weight option for text-to-music workflows.

Tradeoffs to consider

Flagship TTS models (Fish Audio S2.1 Pro, Gemini Flash TTS) produce more nuanced, characterful output at the cost of latency headroom. Realtime-tier models (Inworld TTS 1.5 Mini, Inworld TTS-2) trade some expressive ceiling for low-latency delivery. For live interaction, the latency win is not a compromise: it is the product. For output that ships (audiobooks, narration, ads), expressiveness wins.

TTS and music generation are separate workloads with separate cost structures. TTS is priced per character or per input token. Music generation is priced per second or per song. Avoid comparing them on cost-per-request alone: a 60-second music track and a 60-second narration clip are very different in value and generation complexity.

ACE-Step is the only open-weight pick in this collection and covers music generation only. For open-weight TTS at scale, Alibaba Qwen3 TTS is worth evaluating. Open music generation (ACE-Step) lands within striking distance of frontier MiniMax quality on most prompts: a strong result for a self-hostable pipeline.

All models here are cleared for commercial use through Runware, but underlying license terms vary. Open-weight models carry their own terms on redistribution and training on outputs. Check the model page before any external commercial launch.

Production notes

TTS models bill by input text size, with the unit varying by provider: MiniMax Speech lists per 1,000 characters, Gemini Flash TTS lists per million tokens, and Fish Audio measures millions of UTF-8 bytes of input text. Music models bill per song or per second of output. Each pick row shows the model's own catalog pricing, so plan your spend against your actual script lengths and generation count.

Full pricing reference

For voice agents and interactive apps, streaming time-to-first-audio is the metric that matters, not total generation time. Fish Audio S2.1 Pro ships realtime streaming; Inworld TTS-2 is designed for turn-level streaming with tone carry-forward across turns. If you start with a batch endpoint and switch to streaming later, expect prompt changes and integration work. Wire streaming in your prototype early.

Streaming guide

Every TTS pick in this collection sits behind a Runware-managed endpoint with SLA, but upstream availability varies across providers. Fish Audio S2.1 Pro and MiniMax Speech 2.8 cover the same expressive voice tier. For music generation, MiniMax Music and ACE-Step cover similar creative territory. One-line routing configuration in the Runware API.

Error handling and fallback routing

Pin the model version in your API calls (fishaudio:s2.1@pro, not just a family name). TTS models iterate fast: voice characteristics, language support, and pricing all shift between versions. When a new release lands you want to eval your scripts against it and opt in deliberately.

API versioning reference

Gemini Flash TTS supports inline audio tags in script text and multi-speaker dialogue; xAI TTS ships speech tags. Inworld TTS-2 and Fish Audio S2.1 Pro support free-form natural language voice direction. These control delivery, emotion, and pacing. Invest time in the control layer early.

Voice direction guide

About this collection

The audio generation landscape has moved faster than any other modality in the last eighteen months. This collection answers a single question we hear from product teams building audio features on Runware: which audio models would you ship with today? It covers TTS (expressive, fast, multi-speaker, non-English), music generation (vocal, instrumental, open-weight), and voice agents. Biased toward models with strong naturalness, predictable latency, transparent pricing, and a license you can ship behind.

Every TTS model in the catalog runs a fixed eval set: ten scripts across narration, scripted dialogue, expressive emotional delivery, and multilingual content spanning at least four language families. Scores weight naturalness and prosody most heavily, then intelligibility and emotional accuracy. Latency is measured on realtime-eligible models under actual streaming conditions. Music picks are evaluated on prompt adherence, structural coherence across genres, and production quality.

Every pick links to its model page on Runware with the full API reference, supported parameters, language lists, and pricing examples. The Runware Learning Center covers cross-model audio topics: text normalization for TTS, streaming integration, voice direction and audio tags, LLM-to-speech pipeline design, and music prompt structure.

Browse the Learning Center

New audio models run the eval set the week they land on Runware. When a new entrant clears the bar against the current champion across naturalness, latency, language coverage, and licensing, it replaces the previous pick. Models that fall out of the curated picks remain in the broader audio catalog and appear in the All view.

Every pick in this collection runs on the Runware inference layer: a unified API with predictable latency, transparent pricing, and a single billing surface across TTS and music providers. Routing through Runware means consistent error handling, SLA-backed uptime, version pinning, and the ability to swap between picks without rewriting your integration.

About Runware