Fish Audio lands $52M seed for expressive AI voice models
What's the deal? Palo Alto-based Fish Audio has raised $52 million in a seed round led by Coreline VenturesDealroom has a profile for this one. Try Dealroom → and Capital TodayDealroom has a profile for this one. Try Dealroom →. Participants included 359 CapitalDealroom has a profile for this one. Try Dealroom →, Parable, Play TimeDealroom has a profile for this one. Try Dealroom →, Alphaist PartnersDealroom has a profile for this one. Try Dealroom →, BayhouseDealroom has a profile for this one. Try Dealroom →, Carya Venture PartnersDealroom has a profile for this one. Try Dealroom →, and HF0Dealroom has a profile for this one. Try Dealroom →.
What's the endgame? The startup builds AI voice models for both creators and enterprises, offering more than 15,000 natural language controls. Since launching last year, it has drawn more than 8 million users across its open-source and hosted models.
The numbers: Fish Audio now generates $21 million in annual recurring revenue. Paid customers include HeyGen, Sanas, and Plaud, which use its enterprise APIs and platform.
The product: The company has launched five models in the past year — four for speech generation and one for speech-to-text. Three speech models are open-source, but the latest S2.1 Pro is available only through the paid API. It plans to release an audio understanding model this year and is building a speech-to-speech model.
Fish Audio began as a project by former NVIDIA researcher Shijia LiaoDealroom has a profile for this one. Try Dealroom →, who trained a voice model on a single GPU and open-sourced it. The Fish Speech GitHub repository now has more than 31,000 stars.
Why now? Chief executive officer and co-founder Rissa CaoDealroom has a profile for this one. Try Dealroom → said the open-source project ran efficiently without capital, but the company wanted to build advanced models and serve enterprises as investor interest grew.
What could go wrong? Fish Audio builds its voice library partly by asking users to submit and monetise their own voices. Some creators alleged their voices were uploaded without consent, and takedowns were slow. Cao said the process is now automated, removing flagged voices in under three minutes once a user proves ownership.
"Every enterprise has different use cases and different preferences," Cao said, noting that avatar firms want realism, gaming studios want expressiveness, and voice agents want low latency.
The signal: The speech generation market is crowded, with ElevenLabs, WellSaid, Cartesia, SpeechifyDealroom has a profile for this one. Try Dealroom →, Async, and Krisp all chasing creators and enterprises. Rico Mallozzi, a partner at 359 Capital, said fine-grained developer controls and cost-efficient training will help Fish Audio compete against larger AI labs.
Read more: TechCrunch