Voice AI startup Fish Audio has closed a $52 million seed round — one of the largest seed financings in the AI audio space — just a year after launching as a side project. The company now serves 8 million users, generates $21 million in annual recurring revenue, and counts HeyGen, OpenAI, and LiveKit among its enterprise customers.
Key Highlights
- $52M seed round led by Coreline Ventures and Capital Today, with 359 Capital, 645 Ventures, HF0, and others participating
- $21M ARR and 8M+ users achieved within the company's first year of operation
- S2.1 Pro model clones any voice from a 5-second audio clip in roughly 15 seconds
- 83 languages supported with 15,000+ natural language emotion controls
- 31,000+ GitHub stars for Fish Speech, the open-source speech synthesis project
The S2.1 Pro Model
Fish Audio's flagship offering, S2.1 Pro, is designed to set a new bar for expressive, natural-sounding synthetic speech. The model can clone a voice from as little as five seconds of reference audio and produce output within 15 seconds, making real-time voice personalization practical for the first time at scale.
In independent blind listening tests, S2.1 Pro was preferred by 67% of listeners over competing models. The gap is attributed to the model's granular emotion controls — more than 15,000 natural language instructions that let developers tune expressiveness, pacing, and tone without custom training. The model supports 83 languages and offers on-premises deployment with HIPAA compliance, opening the door to healthcare, legal, and financial use cases.
Currently available exclusively through Fish Audio's paid API, the company plans to make S2.1 Pro free to every developer through its official API starting at the end of August 2026.
Enterprise Traction
The company's customer list reads like a who's who of the voice-AI stack. HeyGen uses Fish Audio's models to power talking-avatar video generation. Retell AI and LiveKit integrate Fish Audio for low-latency voice agents. Telnyx and OpenArt round out a customer base that spans media, robotics, and enterprise communications.
CEO and co-founder Rissa Cao frames the product opportunity in terms of use-case diversity: "Every enterprise has different use cases — a gaming studio would want expressive voices for their characters," while communications companies need "low-latency voices that are expressive enough for calls." Fish Audio's multi-model lineup — four speech-generation models (three open-source) plus one speech-to-text model — is designed to serve that range.
From Frustration to $52M
The company traces its origins to co-founder Shijia Liao's frustration with the synthetic, robotic quality of existing voice models. Liao, a former NVIDIA researcher, teamed with Cao to build a model that sounded human — starting as an open-source project before commercial momentum took over.
The resulting Fish Speech repository has accumulated more than 31,000 GitHub stars, an unusually strong open-source signal for a company at this stage. That community pull drove early adoption and gave the team real-world feedback across languages and use cases before the enterprise product even launched.
What's Next
The $52 million will fund three priorities: expanding the model lineup with voice-native large language models and speech-to-speech tools, deepening integrations with developer platforms like LiveKit and Retell AI, and building out an enterprise sales team to convert the company's existing inbound demand into long-term contracts.
The roadmap signals a broader ambition: Fish Audio wants to own the full audio intelligence stack, not just text-to-speech. A planned audio-understanding model and a speech-to-speech model would let it compete in a market where OpenAI, ElevenLabs, and Cartesia are all racing to define the voice layer of AI applications.
"Voice is becoming the primary interface for AI," Cao said. "Fish Audio has established a strong record by advancing performance, multilingual capabilities, emotional nuance, and affordability."
Source: TechCrunch