Fish Audio Raises $52M for Expressive AI Voice Models

Fish Audio Raises $52M for Expressive AI Voice Models

Fish Audio Secures $52M Seed Funding for Expressive, Steerable AI Voices

Palo Alto-based Fish Audio has raised $52 million in a seed round led by Coreline Ventures and Capital Today to scale its library of more than 15,000 natural language controls for AI voice models. The startup, founded by former Nvidia researcher Shijia Liao, now serves over 8 million users through open-source and hosted models, generating $21 million in annual recurring revenue.

Meeting Diverse Voice AI Needs

Fish Audio targets both creative and enterprise use cases. Its models power AI avatars (e.g., HeyGen), gaming characters, and voice agents (e.g., LiveKit). The company offers five models—four for speech generation and one for speech-to-text—with three open-sourced and the latest S2.1 Pro available via paid API.

Key Features for Developers and Teams

  • Expressive, low-latency voices for real-time applications
  • Voice cloning and fine-grained controls
  • Monthly plans for creators and enterprise APIs

Addressing Voice Ownership and Consent

Fish Audio compensates users who submit voices for training. However, it faced backlash when creators alleged unauthorized uploads. The startup has now automated its DMCA takedown process, enabling voice removal in under three minutes upon proof of ownership.

Industry Call for Verified Voice Ownership

Investor Osuke Honda of Coreline Ventures stresses that trust is critical: “Consent, transparency, and attribution must be built into the product.” He advocates for verified voice ownership, clear licensing, and revenue-sharing models.

Competitive Landscape and Future Plans

The speech-generation market includes ElevenLabs, WellSaid, and Cartesia. Fish Audio differentiates with cost-efficient training and fine-grained controls. The startup plans to release an audio understanding model and a speech-to-speech model this year.

Investor Confidence in Technical Acumen

Rico Mallozzi of 359 Capital praises Fish Audio’s state-of-the-art models: “They’ve closed the gap between artificial-sounding and human-like voices with a small team.” The funding will support advanced model development and enterprise expansion.

AI voice models  text to speech  voice AI startup  Fish Audio  speech generation 

Comment