Fish Audio Just Raised $52M at Seed Stage — and It Already Has $21M ARR to Show for It
Fish Audio closed a $52M seed round led by Coreline Ventures and Capital Today — one year after launching, with $21M in ARR and 8 million users already on the platform.

One year ago, Shijia Liao was a researcher at Nvidia, frustrated by how flat and robotic AI-generated voices sounded. So he trained a speech model on a single GPU and pushed it to GitHub. That repository, Fish Speech, now has more than 31,000 stars. The company that grew out of it, Fish Audio, just closed a $52 million seed round — with $21 million in annual recurring revenue already on the books.
The round was co-led by Coreline Ventures and Capital Today, with 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners, and Bayhouse Ventures also participating. No valuation was disclosed.
From a Side Project to 8 Million Users
Fish Audio, operating legally as Hanabi AI Inc. and headquartered in Palo Alto, launched its hosted platform roughly twelve months ago. Since then it has shipped five models — four for speech generation, one for speech-to-text — and open-sourced three of them. The fourth, S2.1 Pro, is paywalled behind the company’s API.
That model can clone a voice from a five-second audio clip in about 15 seconds. It supports 83 languages and is steerable through more than 15,000 natural language prompts — think tags like [whispers] or [laughing nervously] embedded directly in text input to control delivery at the word level. In blind listening tests, 67% of listeners preferred S2.1 Pro over competing outputs, according to company-reported figures.
Enterprise customers already running on Fish Audio include OpenAI, HeyGen, LiveKit, Retell, Telnyx, and Sanas. Each has a different requirement. HeyGen needs realism for AI avatars. Gaming studios want expressive character voices. Voice agent companies like LiveKit want low latency and enough expressiveness to carry a phone call.
“Every enterprise has different use cases and different preferences,” CEO and co-founder Rissa Cao told TechCrunch. “For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices.”
Why Fish Audio Didn’t Need This Money — Until It Did
Cao told TechCrunch that Fish Audio was running efficiently on open-source distribution and creator subscriptions before the raise. The company didn’t need outside capital, she said. But two things changed: the ambition to build more advanced models, and rising enterprise demand that required dedicated sales infrastructure.
The capital goes toward three initiatives: voice-native large language models, real-time speech-to-speech translation, and an enterprise sales team. An audio understanding model is also planned for release before year-end.
Rico Mallozzi, a partner at 359 Capital, framed the technical bar plainly. “What they’ve been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible,” he told TechCrunch. “It shows their technical acumen in closing the gap between artificial-sounding and human-like voices.”
The Voice Consent Problem
Fish Audio’s community model — paying creators when their submitted voices are used to train models — ran into trouble earlier this year. Creators alleged their voices had been uploaded without consent. The company had a DMCA takedown process, but it was slow.
Cao says that process is now automated. A creator submits a voice sample or contract, and the infringing upload comes down in under three minutes. But the window before a complaint is filed remains open, and voices can continue to circulate on the platform until someone acts.
Oskue Honda, a partner at Coreline Ventures, acknowledged the structural issue directly. “A community-centric approach can only become a durable advantage if creators trust the platform,” Honda said. “That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially.”
That’s not just a values statement — it’s a business risk. The free S2.1 Pro API window closes at the end of August 2026. Fish Audio needs the developer community it attracted with open weights to convert into paying enterprise customers. Anything that erodes creator trust erodes that funnel.
Going After ElevenLabs
The voice synthesis market is genuinely crowded. ElevenLabs raised $500 million at an $11 billion valuation earlier this year. WellSaid, Cartesia, Speechify, and Krisp are all competing for overlapping budgets. Fish Audio’s pitch is a combination of open-source credibility, a creator marketplace, and cost efficiency that larger labs can’t easily replicate at the same price point.
The $52 million seed itself signals something about investor appetite. Entire, the developer platform from former GitHub CEO Thomas Dohmke, pulled in $60M earlier this month. Seed rounds at this scale are no longer exceptional in AI infrastructure — but they are still a bet that the revenue is real and the retention holds.
For Fish Audio, the next six months are the proof. Convert the free-tier developers before August’s API window closes, close enterprise contracts at scale, and get an audio understanding model into production. The $21M ARR gets investors in the door. What comes next determines whether this stays a challenger or becomes the category.
Fish Audio’s voice AI timing is pointed — OpenAI just expanded GPT Live voice to enterprise and education plans, sharpening the competitive stakes across the entire segment.
Frequently Asked Questions
How much did Fish Audio raise and who led the round?
Fish Audio raised $52 million in a seed round co-led by Coreline Ventures and Capital Today, with participation from 359 Capital, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners, Bayhouse Ventures, and Play Time. No valuation was disclosed.
What does Fish Audio’s S2.1 Pro model do?
S2.1 Pro is Fish Audio’s flagship text-to-speech and voice cloning model. It can clone a voice from a five-second clip in roughly 15 seconds, supports 83 languages, and accepts over 15,000 natural language emotion and delivery tags for word-level control. It is available through Fish Audio’s paid API only.
Who are Fish Audio’s enterprise customers?
Named enterprise customers include OpenAI, HeyGen, LiveKit, Retell, Telnyx, and Sanas, each using Fish Audio’s models for different applications ranging from AI avatars to voice agents and real-time call systems.
How does Fish Audio compare to ElevenLabs?
ElevenLabs raised $500 million at an $11 billion valuation earlier in 2026, making it the current market leader by capital and valuation. Fish Audio’s competitive pitch centres on open-source roots, a creator voice marketplace, and a lower-cost API — particularly attractive to developers and mid-market enterprises that don’t need ElevenLabs’ scale or pricing tier.