-
Palo Alto-based Fish Audio has raised $52 million in a seed round led by Coreline Ventures and Capital Today to scale its library of more than 15,000 natural language controls for AI voice models. The startup, founded by former Nvidia researcher Shijia Liao, now serves over 8 million users through open-source and hosted models, generating $21 million in annual recurring revenue.
Fish Audio targets both creative and enterprise use cases. Its models power AI avatars (e.g., HeyGen), gaming characters, and voice agents (e.g., LiveKit). The company offers five models—four for speech generation and one for speech-to-text—with three open-sourced and the latest S2.1 Pro available via paid API.
Fish Audio compensates users who submit voices for training. However, it faced backlash when creators alleged unauthorized uploads. The startup has now automated its DMCA takedown process, enabling voice removal in under three minutes upon proof of ownership.
Investor Osuke Honda of Coreline Ventures stresses that trust is critical: “Consent, transparency, and attribution must be built into the product.” He advocates for verified voice ownership, clear licensing, and revenue-sharing models.
The speech-generation market includes ElevenLabs, WellSaid, and Cartesia. Fish Audio differentiates with cost-efficient training and fine-grained controls. The startup plans to release an audio understanding model and a speech-to-speech model this year.
Rico Mallozzi of 359 Capital praises Fish Audio’s state-of-the-art models: “They’ve closed the gap between artificial-sounding and human-like voices with a small team.” The funding will support advanced model development and enterprise expansion.
Comment