How well do AI mental health companions actually work?
The best evidence comes from a 2025 meta-analysis of 14 randomized controlled trials (RCTs) involving 6,314 people. It found that generative AI chatbots reduce negative mental health symptoms like depression and anxiety by a small but statistically significant amount (effect size 0.30) [1]. That means the average person using one does better than about 62% of people who don't. However, the 95% prediction interval ranged from -0.85 to 1.67, which means in some real-world settings the chatbot could actually make things worse [1]. This wide range shows the average effect hides a lot of uncertainty.
The same review found that chatbots designed for social interaction (like a friend) were more effective than task-oriented ones (like a symptom tracker) [1]. This suggests that the quality of the interaction matters more than just having an AI present. A separate 2025 experiment found that when AI chatbots used highly personalized, warm messages, users felt more emotionally validated — and this effect was stronger when users felt a sense of social presence with the bot [3]. So the evidence points to 'how' the AI talks being critical, but we don't yet know which specific conversation styles work best for different people.
What are the biggest safety and ethical gaps?
Safety is the most glaring gap. A 2025 scoping review of 101 articles on conversational AI in mental health found that safety and harm were discussed in only 52 of them (51.5%) [2]. The top concerns were handling suicidal users, giving harmful or wrong suggestions, and users becoming dependent on the AI. That means nearly half the published work doesn't even address basic safety issues. The same review found that privacy and confidentiality were discussed in 62 articles (61.4%), but accountability — who is responsible if the AI gives bad advice — was only covered in 31 (30.7%) [2]. For a tool that could be used by vulnerable people, these are huge gaps.
Another gap is who has been studied. The 2025 meta-analysis found that most generative AI chatbot studies were done in non-WEIRD countries (non-Western, Educated, Industrialized, Rich, Democratic), and almost none focused on young children or older adults [1]. This means we don't know if these tools work safely for the very young or the elderly. A 2024 survey of 800 users of the Sakhi chatbot found that female users aged 30 and above were the majority, and factors like age and work changes predicted stress [5]. But the machine learning models used to analyze that data had low accuracy (Random Forest 0.37, Gradient Boosting 0.33), meaning the predictions were barely better than chance [5]. So even the data we have from real users is noisy and unreliable.
What evidence is missing for long-term use?
There is almost no long-term evidence. The 2025 meta-analysis included only 14 RCTs, and the authors noted this small number as a major limitation [1]. None of the studies tracked users for more than a few months, so we don't know if benefits last, if users become dependent, or if the AI's advice becomes less accurate over time. The Jarvie system, described in a 2025 paper, claims to 'continuously adapt through user interactions' [4], but no long-term data on safety or effectiveness was provided. Similarly, the Sakhi chatbot offers 'remote mental health monitoring' [5], but without evidence on what happens when monitoring detects a crisis.
The ethical review highlighted that only 12 out of 101 articles (11.9%) discussed user autonomy — meaning the user's ability to control or stop the interaction [2]. If a companion AI is designed to keep users engaged for long periods, that could be helpful or harmful, but we have no data to tell which. The review also found that only 16 articles (15.8%) discussed concerns about healthcare workers' jobs [2], which means we don't know how these tools should fit into existing care systems. Without long-term studies, we can't say whether AI companions are a safe supplement or a risky substitute for human care.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2024 to 2025, 5 from 2024 or later, 3 in Q1 journals, collectively cited 67 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 39 papers retrieved from a database of over 500 million.
Sources used in this answer
Generative AI Mental Health Chatbots as Therapeutic Tools: Systematic Review and Meta-Analysis of Their Role in Reducing Mental Health Issues
A 2025 meta-analysis of 14 RCTs (6,314 participants) found generative AI chatbots reduce negative mental health symptoms with an effect size of 0.30, but the prediction interval (-0.85 to 1.67) suggests some users may worsen; most studies were in non-WEIRD countries and lacked data on children and older adults.
Exploring the Ethical Challenges of Conversational AI in Mental Health Care: Scoping Review
A 2025 scoping review of 101 articles identified 10 ethical themes, with safety and harm discussed in only 52% of articles, privacy in 61%, and accountability in 31%; crisis management and harmful suggestions were top concerns.
Artificial intelligence chatbots as a source of virtual social support: Implications for loneliness and anxiety management
A 2025 experiment showed that AI chatbots using highly person-centered (warm, personalized) messages increased users' emotional validation, and this effect was stronger when users felt social presence with the bot.
Jarvie: AI-Driven Mental Health Companion
A 2025 paper describes 'Jarvie,' an AI companion using NLP and sentiment analysis to provide real-time emotional support, but provides no long-term safety or effectiveness data.
Sakhi: AI-Generated Mental Health Companion
A 2024 study of the Sakhi chatbot surveyed 800 users (mostly women aged 30+) and found age and work changes predicted stress, but machine learning models had low accuracy (Random Forest 0.37, Gradient Boosting 0.33).
