Review Maps 21 Studies of Generative AI Mental Health Chatbots

A scoping review published July 23 examined 21 studies of generative AI mental health chatbots across 11 countries. Users often reported convenience, personalization and empathy, but most interventions were early-stage and engagement commonly declined over time, leaving limited evidence on clinical efficacy, long-term use and standardized safety evaluation.
A scoping review published in npj Digital Medicine on July 23, 2026, maps how purpose-built generative AI mental health chatbots have been designed and experienced by users. The researchers searched seven databases, screened 1,899 articles and included 21 studies conducted across 11 countries from 2023 through 2025.
Most of the interventions were early-stage and based on cognitive behavioral therapy principles. They most often targeted depression and anxiety, while other studies addressed stress, loneliness, eating disorders, post-traumatic stress disorder and psychological concerns associated with dementia. Sample sizes ranged from five to 527 participants, and the reviewed tools ranged from prototypes to clinical trials and one real-world implementation.
User experience was promising but uneven
The included studies generally reported moderate-to-high usability, therapeutic alliance and user satisfaction. Participants often valued convenience, personalized responses and perceived empathy. About two-thirds of the interventions used non-embodied text chatbots, while others added voice, avatars, augmented reality or other multimodal features.
Those positive impressions did not consistently translate into stronger clinical outcomes or sustained use. The review found that engagement often declined during multi-week interventions. Users also reported repetitive or contextually mismatched responses, privacy concerns, reduced human contact and uncertainty about how systems would respond in a crisis.
The authors describe features such as domain grounding, adaptive tailoring, structured delivery and co-design as associated with better reported experiences. They do not establish that those features caused better outcomes, because the studies used heterogeneous measures and offered limited direct comparisons.
What the evidence can and cannot show
This is a scoping review of intervention design and user experience, not a pooled estimate of clinical effectiveness. It identifies gaps in the evidence base rather than proving that generative AI chatbots are effective or ineffective mental health treatments.
For teams developing conversational health systems, the practical implication is that fluent and empathetic interaction is only one layer of evaluation. Longitudinal outcome measurement, crisis-response testing, privacy safeguards, transparent reporting and independent assessment remain necessary before user satisfaction can be treated as evidence of dependable clinical value.
Key Points
- 1The review included 21 studies across 11 countries after screening 1,899 articles from seven databases.
- 2Users often reported convenience, personalization and empathy, but engagement commonly declined and outcome measures varied widely.
- 3The authors call for standardized user-experience and safety evaluation, efficacy trials, transparent reporting and inclusive co-design.
Scoring Rationale
The peer-reviewed scoping review provides a structured map of 21 studies in a high-stakes application area and identifies concrete evaluation gaps. Its practical value is limited by heterogeneous, mostly early-stage evidence and the absence of pooled clinical-effect estimates.
Sources
Primary source and supporting public references used for this report.
Practice with real Health & Insurance data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Health & Insurance problems
