Skip to main content

Fish Audio Anniversary: S2.1Pro Model Brings Real-Time Speech with Emotion Control

Fish Audio is marking its anniversary with the release of S2.1Pro, its most advanced production-level speech model yet. Unlike traditional text-to-speech systems that stick to scripted narration, this one is built for real-time conversation. It's now available through the Fish Audio API, and there's a free version with reasonable limits for developers to test and tinker.

Image

Performance and Features

On the technical side, S2.1Pro delivers first-frame audio playback in about 90 milliseconds—fast enough for natural, smooth back-and-forth dialogue. It supports 83 languages under a unified speech recognition architecture. But what really sets it apart is the flexibility in voice control. Forget pre-set emotion menus; you can embed free-form bracket tags directly in the text to trigger whispers, tense laughter, or other fine-grained instructions on the fly.

The model also natively handles multi-speaker dialogue generation. Need to clone a voice? Just 10 to 30 seconds of reference audio is enough to replicate the target tone and speaking style accurately, without any extra fine-tuning.

Implications

S2.1Pro signals a shift in speech AI—from rigid, script-based synthesis toward real-time, high-fidelity interaction. It lowers the barrier for developers and provides a solid computing foundation for deploying intelligent voice applications worldwide.

Key Points

  • First-frame latency of ~90ms for real-time dialogue
  • Supports 83 languages with unified speech recognition
  • Emotion control via inline bracket tags (e.g., whispers, laughter)
  • Multi-speaker generation and voice cloning from 10-30 seconds of audio
  • Available via Fish Audio API with a free tier for testing