Where is the evidence strongest? Mental health chatbots have multiple randomized trials showing real symptom reduction.
The most rigorous prospective clinical evidence for AI triage chatbots comes from mental health, where several randomized controlled trials (RCTs) have been conducted. In a 16-week RCT with 83 university students, a chatbot-delivered depression intervention significantly reduced depression scores (PHQ-9) and anxiety scores (GAD-7) compared to a bibliotherapy control group [1]. The effect was substantial: the chatbot group showed a statistically significant reduction in depression (F=22.89, p<0.01) and anxiety (F=5.37, p=0.02), and participants reported higher therapeutic alliance with the chatbot than with the self-help book [1].
A second RCT with 84 college students found that chatbots with high-social-cue designs (text + voice + animations) outperformed text-only versions, producing greater reductions in depression (effect size d=0.63, p<0.01) and anxiety (d=0.50, p=0.003), along with higher adherence, satisfaction, and therapeutic alliance [4]. Most recently, a national RCT of 210 adults with clinically significant depression, anxiety, or eating disorder symptoms tested a generative AI chatbot (Therabot) against a waitlist control and found significant symptom improvement at 4 and 8 weeks [7]. These three RCTs [1][4][7] converge on the same conclusion: AI chatbots can be effective mental health interventions, with the largest and most recent trial [7] providing the strongest evidence.
For radiology and emergency triage, the evidence is promising but mixed — AI improves speed but not always accuracy.
In radiology, prospective studies show AI triage can reduce wait times but has not consistently improved diagnostic accuracy. A prospective single-center study of an AI triage system for pulmonary embolism on CT scans found that mean wait time for positive cases dropped from 21.5 minutes without AI to 11.3 minutes with AI (p<0.001), but radiologist accuracy (97.6% vs. 98.6%) and miss rate (12.3% vs. 6.1%) did not significantly improve [2]. In contrast, a real-world evaluation of an AI triage system for chest X-rays across 43 radiologists found high sensitivity (82-93%) and specificity (91-99%) for urgent, non-urgent, and normal categories, and significantly reduced turnaround times across all subgroups [3]. The difference may reflect the specific task: the chest X-ray study used a 3-tier triage system [3], while the pulmonary embolism study focused on binary detection [2].
For emergency triage, a cross-sectional study of 46 emergency case scenarios found that ChatGPT-4o showed moderate-to-substantial agreement with expert physicians (kappa=0.695) and had 100% sensitivity for the most critical (Level 1) and least urgent (Level 5) cases, but only 50% sensitivity for intermediate urgency (Level 4) [9]. The authors concluded that AI chatbots are not yet ready to replace clinicians but could serve as decision-support tools [9]. A separate study of ophthalmic triage found that ChatGPT (GPT-4) listed the correct diagnosis among its top three suggestions in 93% of 44 clinical vignettes and gave appropriate triage urgency in 98% of cases — comparable to physician respondents [6]. However, another chatbot (Bing Chat) performed worse, with 77% diagnostic accuracy and a tendency to overestimate urgency [6], highlighting that performance varies greatly by model.
Real-world feasibility and acceptance are high, but data gaps remain for many clinical settings.
Beyond efficacy trials, prospective clinical studies have demonstrated that AI triage chatbots are feasible and acceptable in real-world settings. A study at a cancer center tested an AI chatbot for hereditary breast and ovarian cancer screening: all 11 participants completed the chatbot interaction, and its determinations matched those of certified genetic counselors [5]. The chatbot also provided new family history information for 27% of participants that the counselors had not obtained [5]. A survey of 677 medical staff in China found 77.1% acceptance of AI triage, with 45.2% preferring AI triage exclusively [8]. Direct experience with AI was the strongest predictor of preference [8].
However, the evidence base has important gaps. Most studies are small (11-210 participants) [1][4][5][7], many are single-center [2][3][5][6], and follow-up periods are short (4-16 weeks) [1][4][7]. No study here prospectively compared AI triage chatbots against standard clinical triage across a broad range of emergency conditions in a real emergency department. The strongest evidence is concentrated in mental health [1][4][7] and specific radiology tasks [2][3], while general emergency triage relies on simulated vignettes [6][9]. As one study noted, AI chatbots are 'not yet ready to replace clinical professionals' but can serve as effective decision-support tools [9].
About These Sources
This answer is built on 9 peer-reviewed studies — published from 2022 to 2025, 5 from 2024 or later, 4 in Q1 journals, collectively cited 517 times — selected as the most relevant from 9 studies that passed quality screening, drawn from 61 papers retrieved from a database of over 500 million.
Sources used in this answer
Using AI chatbots to provide self-help depression interventions for university students: A randomized trial of effectiveness
In a 16-week RCT with 83 university students, a chatbot-delivered depression intervention significantly reduced depression (PHQ-9) and anxiety (GAD-7) scores compared to bibliotherapy, and participants reported higher therapeutic alliance with the chatbot.
Prospective Evaluation of AI Triage of Pulmonary Emboli on CT Pulmonary Angiograms
In a prospective single-center study of 1,526 CT pulmonary angiograms, an AI triage system reduced mean wait time for positive pulmonary embolism cases from 21.5 to 11.3 minutes but did not significantly improve radiologist accuracy or miss rate.
Real-World evaluation of an AI triaging system for chest X-rays: A prospective clinical study
In a real-world evaluation of an AI triage system for chest X-rays across 43 radiologists, the system showed high sensitivity (82-93%) and specificity (91-99%) for urgent, non-urgent, and normal categories, and significantly reduced turnaround times across all subgroups.
Depression intervention using AI chatbots with social cues: a randomized trial of effectiveness.
In a 16-week RCT with 84 college students, chatbots with high-social-cue designs (text+voice+animations) produced greater reductions in depression and anxiety, and higher adherence, satisfaction, and therapeutic alliance compared to text-only chatbots.
Preliminary Screening for Hereditary Breast and Ovarian Cancer Using an AI Chatbot as a Genetic Counselor: Clinical Study.
In a clinical feasibility study at a cancer center with 11 participants, an AI chatbot for hereditary breast and ovarian cancer screening correctly matched genetic counselors' determinations and provided new family history information for 27% of participants.
Artificial intelligence chatbot performance in triage of ophthalmic conditions
In a cross-sectional study of 44 ophthalmic vignettes, ChatGPT (GPT-4) listed the correct diagnosis among its top three in 93% of cases and gave appropriate triage urgency in 98%, comparable to physician respondents; Bing Chat performed worse.
Randomized Trial of a Generative AI Chatbot for Mental Health Treatment
In a national RCT of 210 adults with clinically significant depression, anxiety, or eating disorder symptoms, a generative AI chatbot (Therabot) significantly improved symptoms at 4 and 8 weeks compared to a waitlist control.
AI triage or manual triage? Exploring medical staffs’ preference for AI triage in China
A survey of 677 medical staff in China found 77.1% acceptance of AI triage, with 45.2% preferring AI triage exclusively; direct experience with AI was the strongest predictor of preference.
Evaluating the Accuracy of Artificial Intelligence Chatbots in Triaging Emergency Cases: A Comparative Study with Expert Clinicians
In a cross-sectional study of 46 emergency case scenarios, ChatGPT-4o showed moderate-to-substantial agreement with expert physicians (kappa=0.695) and 100% sensitivity for the most critical cases, but only 50% sensitivity for intermediate urgency cases.
