Fish Audio secures $52 million to offer voice AI for free and monetize enterprises.
Fish Audio, a voice-AI startup based in Palo Alto, has secured $52 million in what it terms a seed funding round, as initially reported by TechCrunch. The company has an unusual business model: it provides its top models for free while charging for the associated infrastructure.
The startup, which has been operating for a year, claims that over 8 million people are using its models through either open-weight releases or its hosted platform, generating $21 million in annual recurring revenue.
The strategy of giving away models and charging for performance is a departure from conventional approaches, leveraging an open-source strategy applied to voice technology. In just one year, Fish Audio has released five models, three of which have been open-sourced. Offering free weights along with a free frontier model helps them gain traction among developers, while revenue is generated through contracts with enterprises.
Even its latest model, S2.1 Pro, which is not openly released, remains free via the API until August 31. This model can replicate a voice from a five-second audio clip, supports 83 languages, and provides audio output in approximately 70 milliseconds. Paid plans include latency and uptime guarantees, and Fish requests that companies with revenues over $1 million consult before using the free tier, as reported by Unite.AI.
In summary, the business relies on attracting developers with open weights while enterprises pay for reliability assurances. Companies such as HeyGen, LiveKit, Retell, Sanas, and OpenArt already utilize its APIs. The $52 million “seed” investment at $21 million in revenue illustrates the rapid evolution of voice technology from being a mere feature to a vital component of infrastructure.
Voice is becoming the primary interface for interacting with AI, increasingly becoming the standard for communication, whether in systems like Claude’s voice mode or in entities managing sales and support calls. CEO Rissa Cao notes that the demand varies by application—avatar companies seek realism, game developers aim for expressive characters, and voice-agent enterprises prioritize low latency while maintaining a natural sound.
However, the landscape is competitive and well-capitalized, with ElevenLabs valued at approximately $22 billion and significant investment flowing into competitors such as Bland. Fish Audio argues that its unique inference stack operates at a significantly lower cost, claiming that S2.1 Pro is cheaper to run compared to ElevenLabs.
The company’s voice library, which is sourced from users who submit their own voices for a fee when utilized, serves as both an asset and a vulnerability. They currently have over two million voices in their library.
This very library has been the subject of controversy; earlier in the year, creators reported that their voices had been uploaded without permission, and the process for removing unauthorized uploads was perceived as slow. In response, Fish has established a dedicated dispute process and claims that removals are now completed in under three minutes. However, this solution only addresses issues after they arise, as there is nothing preventing someone from uploading a voice without consent, leaving it in use until the original owner notices and takes action.
Fish’s principal investor highlighted this issue, stating, “A community-centric approach can only become a durable advantage if creators trust the platform,” according to Osuke Honda from Coreline Ventures, who co-led the funding round. He emphasized that consent, transparency, and attribution must be integral to the product rather than being afterthoughts.
Thus, the funds raised represent a gamble where both the advantage and the risk are intertwined, reflecting the broader discussion surrounding open-weight systems and consent.
Other articles
Fish Audio secures $52 million to offer voice AI for free and monetize enterprises.
Fish Audio secured $52 million for its voice AI, which it provides at no charge while charging businesses for enhanced speed and uptime. Its library of 2 million voices serves as both a competitive advantage and a potential risk.
