Palo Alto-based AI voice startup Fish Audio has secured $50 million in a seed funding round, the company announced Tuesday, as it looks to scale its platform of expressive, steerable voice models for both individual creators and large enterprises. The round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
From a single GPU project to $21M ARR
Fish Audio began as a side project from former NVIDIA researcher Shijia Liao, who grew frustrated with the limited expressiveness of synthetic voices available on the market. Training a voice generation model on a single GPU, Liao open-sourced the result as Fish Speech, a repository that has since accumulated more than 31,000 stars on GitHub. The project quickly gained traction among indie developers, video game designers, and content creators looking for more natural-sounding AI voices.
Also read: US Grid Operator Will Cut Power to Large Data Centers to Prevent Blackouts
Since its formal launch last year, the startup has released five models: four for speech generation and one for speech-to-text. Three of its speech generation models remain open-source, while its latest S2.1 Pro model is available exclusively through a paid API. The company now serves more than 8 million users across its open-source and hosted versions and generates $21 million in annual recurring revenue.
Targeting creators and enterprises with fine-grained controls
Fish Audio offers paid monthly plans tailored to creators and teams, unlocking a set number of generation minutes along with voice cloning features. An enterprise tier provides API access and a dedicated platform, with organizations like HeyGen, Sanas, and Plaud already using the technology.
“Every enterprise has different use cases and different preferences,” Fish Audio CEO and co-founder Rissa Cao told TechCrunch. “For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls.”
The startup’s library of over 15,000 natural language controls allows developers to fine-tune pitch, tone, pacing, and emotional inflection, a capability that Cao says distinguishes Fish Audio from competitors that offer more rigid voice outputs.
Managing consent and trust in a community-driven model
One of Fish Audio’s methods for building its voice library has been to invite users to submit their own voices for training, compensating them when their voices are used commercially. That approach ran into trouble several months ago when some creators alleged their voices had been uploaded without consent. The startup had a DMCA takedown process in place, but removals took a long time.
Cao said the company has now automated the process. Creators can submit a short voice sample or a contract to prove ownership, and their voice will be removed from the platform in less than three minutes. Still, the system remains reactive: an artist’s voice can be uploaded without their knowledge, and it stays on the platform until they file a takedown request.
Oskue Honda, a partner at Coreline Ventures, emphasized that a community-driven model depends on trust. “A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts,” he said. “I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially.”
Competing in a crowded market with efficiency
The speech generation market is increasingly crowded, with well-funded players like ElevenLabs, WellSaid, Cartesia, Speechify, and Krisp all vying for creators and enterprise budgets. Fish Audio’s ability to develop state-of-the-art models with a relatively small team is a key differentiator, according to Rico Mallozzi, a partner at 359 Capital.
“I think what they’ve been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible,” Mallozzi said. “It shows their technical acumen in closing the gap between artificial-sounding and human-like voices.”
Cao noted that the company had been running efficiently without outside capital when it was solely focused on its open-source project and creator plans. The decision to seek funding came as the team wanted to develop more advanced models and accommodate growing enterprise demand alongside rising investor interest.
Looking ahead, Fish Audio plans to release an audio understanding model later this year and is also building a speech-to-speech model. The company’s ability to offer fine-grained controls for developers while keeping model training cost-efficient will be critical to competing against larger AI labs with deeper pockets.
Frequently Asked Questions
What is Fish Audio?
Fish Audio is a Palo Alto-based startup that develops AI voice generation models, offering both open-source and paid API versions. It has over 8 million users and a library of 15,000+ voice controls.
How much funding did Fish Audio raise?
Fish Audio raised $50 million in a seed round led by Coreline Ventures and Capital Today, with participation from several other investors.
Who are Fish Audio’s main competitors?
Fish Audio competes with companies like ElevenLabs, WellSaid, Cartesia, Speechify, and Krisp in the AI voice generation market.
How does Fish Audio handle voice consent issues?
Fish Audio has automated its DMCA takedown process, allowing creators to remove unauthorized voice uploads in under three minutes by submitting a voice sample or contract.

Be the first to comment