Wispr Raises $280M for AI Dictation Platform
Wispr raised $280 million in Series B financing at a $2 billion valuation on August 17, bringing its total funding to $361 million. Menlo Ventures led the round, according to SiliconANGLE and TechCrunch. Alongside the financing, Wispr previewed Canto, a proprietary speech model that the company says reduces Flow's noisy-environment dictation error rate to roughly 5%-10%.
Wispr raised $280 million in Series B financing at a $2 billion valuation, bringing its total capital raised to $361 million. Menlo Ventures led the round, according to SiliconANGLE, with existing backers Notable Capital, NEA, Neo Ventures, 8VC, and MVP Ventures participating alongside new investors including Acrew, Forerunner, Goodwater, Peak XV, Together Fund, and PLUS Capital.
The San Francisco company also previewed Canto, its first proprietary AI model for speech understanding. Wispr's core product, Flow, turns spoken input into formatted text across applications, including by removing filler words and correcting grammar, according to SiliconANGLE.
Canto targets dictation outside quiet rooms
SiliconANGLE reports that Canto was trained using varied speech conditions, including background noise, interruptions, other voices, and environmental sounds. The publication describes this as an effort to improve dictation in settings such as cars, offices, and other noisy environments, where automatic speech recognition commonly degrades.
Wispr told SiliconANGLE that integrating Canto into Flow reduces the error rate for noisy-environment dictation from 30% to approximately 5%-10%. TechCrunch reported the company described the target as reducing errors from 30% to below 10%. Those figures are company-reported performance claims, not independent benchmark results, and neither report specifies the evaluation dataset, language coverage, error metric, or test protocol.
That missing detail matters for ML practitioners assessing speech systems. Word error rate can vary substantially with microphone quality, accents, overlapping speakers, noise type, and whether a system is scored on raw transcription or post-processed text. Speech products that remove disfluencies and rewrite phrasing can be more useful for dictation, but their output cannot be directly compared with a conventional verbatim ASR benchmark without a clearly defined evaluation setup.
Expanding beyond dictation
TechCrunch reports that Wispr recently released a meeting note-taking product, which produces summaries and action items, placing it among products from Granola, Fireflies, and Read AI. The company has also partnered with hardware makers, including Plaud, to enable quieter dictation on their devices, according to TechCrunch.
FinSMEs reports that Wispr intends to use the funding for technical and research hiring, new products, global infrastructure, and its human-computer interaction work. The company is led by CEO Tanay Kothari and serves millions of users across more than 125,000 companies, according to FinSMEs.
The financing arrives amid greater competition in AI dictation, TechCrunch reports, including lower-cost tools developed for prosumers. In comparable speech-product markets, differentiation increasingly depends on robustness under real acoustic conditions, latency, editing quality, privacy controls, and integrations with the systems where users already write and collaborate. Public benchmark methodology and independently reproducible accuracy tests would give engineering teams a clearer basis for evaluating claims in this category.
Key Points
- 1Wispr raised $280 million at a $2 billion valuation, providing substantial capital for an AI voice-interface company competing in dictation software.
- 2Canto's reported 5%-10% noisy-environment error rate targets a persistent speech-recognition failure mode, though public reports lack benchmark and methodology details.
- 3Across comparable dictation products, production evaluation depends on acoustic robustness, latency, privacy, editing behavior, and workflow integrations rather than accuracy alone.
Scoring Rationale
The $280 million Series B and $2 billion valuation make this a notable financing event in voice AI. Canto's claimed noisy-environment performance is relevant to speech ML practitioners, but the available reporting does not provide independent benchmarks or technical evaluation details.
Sources
Public references used for this report.
Practice with real Telecom & ISP data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Telecom & ISP problems


