Fish Audio has positioned S2.1 Pro as its recommended production voice model, built around real-time conversational speech rather than the scripted narration that older text-to-speech systems were designed for. The model is available now through the Fish Audio API, alongside a free tier that runs the same model at no cost for development and testing under fair-use limits. The company presents it as its most capable production model so far, timed to a company anniversary.
The technical pitch centers on latency and language coverage. S2.1 Pro reports a time-to-first-audio of around 90 milliseconds on standard calls, fast enough for natural turn-taking in live dialogue, and it covers 83 languages with a single voice identity that holds across them. Delivery is steered through free-form bracket tags written directly into the text, so a line can carry instructions like a whisper or a nervous laugh without switching to a fixed menu of preset emotions. The model also supports multi-speaker dialogue and voice cloning from short reference samples, typically in the range of 10 to 30 seconds, capturing tone and speaking style without additional fine-tuning.

S2.1 Pro is built as an improvement over the earlier S2-Pro across quality, latency, and throughput, and it carries production guarantees around time-to-first-audio and data processing for teams running it at scale. It plugs into agent workflows through MCP and agent-skill support, which points the model at developers wiring speech into voice agents, phone systems, and long-form audio pipelines. The free tier lowers the barrier for smaller teams and prototyping, offering the same model quality and language coverage as the paid production string without a hard usage cap.
Start testing Fish Audio
Fish Audio builds speech models for creators, developers, and enterprises, with a platform spanning voice generation, voice cloning, and real-time voice applications. Its open-source Fish Speech line has drawn a developer following, passing 20,000 stars on GitHub, and the company has since moved through the OpenAudio S1 generation into the S2 family, which shipped with open weights earlier in the year. The S2 models were trained on large multilingual audio datasets and, by the company's own benchmarks, reached low word error rates against other evaluated systems. S2.1 Pro is the production layer on top of that research, aimed at the growing set of products where voice is becoming the primary interface and latency, language reach, and cloning fidelity decide whether an assistant feels present or delayed.