Fish Audio Raises $52M for Voice AI

Fish Audio, formally Hanabi AI Inc., raised $52 million in seed funding on July 28 for its AI voice platform. SiliconANGLE reports that Coreline Ventures and Capital Today led the round, while Dealroom reports the company serves more than eight million users and generates more than $21 million in annual recurring revenue. The funding places a sizable early-stage bet on expressive text-to-speech, voice cloning, and voice-agent tooling.
Fish Audio, formally Hanabi AI Inc., raised $52 million in seed funding for its AI voice platform. SiliconANGLE reports that Coreline Ventures and Capital Today led the round, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners, and unnamed angel investors.
The company was founded from work by co-founder and chief scientist Shijia Liao, a former Nvidia video researcher, according to SiliconANGLE. The report describes Liao's earlier open-source project, Fish Speech, as having accumulated more than 31,000 GitHub stars.
Voice generation capabilities
Fish Audio provides text-to-speech, voice-cloning, and voice-agent capabilities. SiliconANGLE reports that its platform offers word-level controls for emotion, tone, inflection, and pacing through more than 15,000 natural-language prompts. The company claims that it can clone a voice from a five-second audio sample in less than 15 seconds and supports 83 languages.
Its flagship model is S2.1 Pro. SiliconANGLE reports that, in company-cited listener tests, 67% of listeners preferred Fish Audio outputs over competing systems. That result is a vendor-reported benchmark rather than an independent evaluation, so teams assessing voice quality would need to validate it against their own languages, speaker styles, latency targets, and consent requirements.
Dealroom reports that Fish Audio has more than eight million users and over $21 million in annual recurring revenue. It also reports that the company intends to offer S2.1 Pro free to developers beginning in August and to expand beyond text-to-speech capabilities.
A competitive voice infrastructure market
The round is unusually large for a seed-stage voice AI company. Public reporting frames the funding around an ambition to make voice a default interface for AI models, though Fish Audio's future product scope and commercial terms beyond the reported August offer were not detailed in the available material.
For ML practitioners, expressive speech generation increasingly involves more than intelligible text-to-speech. Comparable systems are evaluated on multilingual pronunciation, speaker consistency, emotion controllability, streaming latency, safety controls around voice replication, and deployment economics. Open-source adoption can supply a meaningful developer distribution channel in this market, while production use generally raises separate questions around provenance, authorization, and abuse prevention.
Key Points
- 1Fish Audio raised $52 million in seed financing, giving voice-generation infrastructure a substantial new early-stage capital infusion.
- 2The platform combines text-to-speech, voice cloning, multilingual support, and word-level expressive controls for developer-facing audio applications.
- 3Comparable voice AI deployments commonly require separate validation of latency, multilingual quality, consent, provenance, and voice-cloning misuse controls.
Scoring Rationale
A $52 million seed round is a major financing event for a developer-focused voice AI company with reported user and revenue traction. The story is relevant to teams building speech interfaces and voice agents, although the available reporting does not establish an independently verified technical breakthrough.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems