WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

Can AI triage chatbots improve patient outcomes in real-world care?

AI triage chatbots can improve patient outcomes in specific settings like stroke care, but real-world evidence is mixed and gaps remain in primary care and equity.

Direct answer

Yes, AI triage chatbots can improve patient outcomes in specific, well-defined clinical scenarios, but the evidence is not yet universal. The strongest data comes from stroke care, where an AI triage system reduced median notification times for neuroendovascular teams from 40 to 25 minutes, a 38% improvement, which is critical because faster treatment directly reduces brain damage [1]. In ophthalmology, ChatGPT matched physician-level diagnostic accuracy (93% vs. 95%) and was actually better at correctly triaging urgency (98% vs. 86%) [2]. However, across the six studies reviewed, the evidence is strongest in emergency and specialty settings, while real-world data from primary care remains scarce and raises concerns about safety and equity [5]. So, while AI chatbots show clear promise in high-stakes, time-sensitive situations, their benefits in broader, everyday care are still unproven.

6sources cited

This article was generated with WisPaper-powered search and paper analysis.

Where does AI triage clearly improve outcomes?

The most compelling evidence for AI triage improving patient outcomes comes from acute stroke care, where every minute of delay can mean permanent brain damage. In a study of 55 patients transferred for endovascular therapy (a procedure to remove a blood clot), implementing the Viz LVO AI triage system cut the median time from a patient's arrival at the initial hospital to notification of the neuroendovascular team from 40 minutes down to 25 minutes — a 38% reduction [1]. This faster notification is clinically meaningful because it allows the specialist team to prepare sooner, and the study also found that the time from first hospital door to the start of the procedure was 25 minutes shorter, though this second difference did not reach statistical significance [1]. The key takeaway: in time-critical emergencies like stroke, AI triage can streamline workflows and reduce delays that directly affect patient survival and recovery.

In ophthalmology (eye care), a 2023 study tested how well AI chatbots triaged 44 common eye complaints compared to human physicians. ChatGPT (using the GPT-4 model) listed the correct diagnosis among its top three suggestions 93% of the time, nearly matching the physicians' 95% [2]. More importantly, ChatGPT correctly identified the urgency of the condition — whether it was an emergency or could wait — in 98% of cases, outperforming the physicians' 86% [2]. This suggests that for patients wondering 'should I go to the emergency room for this eye problem?', an AI chatbot could provide more accurate triage advice than a human doctor in some cases. However, the study used clinical vignettes (written scenarios), not real patients, so real-world performance may differ.

Why isn't this evidence enough to declare AI triage a universal success?

The gap between best-case and typical-case evidence is significant. The strongest studies — like the stroke and ophthalmology ones — were conducted in controlled settings or with specific, high-stakes conditions. In contrast, a 2025 review of AI triage in primary care (general practice) found that current evidence is drawn almost entirely from 'retrospective validations, emergency settings, or vignettes, with scant evaluation of real-world outcomes' [5]. The authors warned that without real-world testing in primary care, AI triage could actually widen health inequalities, because algorithms may perform differently across age, ethnicity, language, and socioeconomic groups [5]. This is a critical caveat: a chatbot that works well for a middle-aged English speaker with a classic stroke might fail for an elderly non-English speaker with atypical symptoms.

Even in settings where AI triage shows promise, real-world deployment reveals challenges. A 2025 study of a peri-operative AI chatbot (PEACH) in a hospital achieved 97.9% accuracy across 240 real clinical interactions, and clinicians said it expedited decisions in 95% of cases [3]. However, the study also noted that the system required iterative updates to fix protocol errors, and its accuracy was measured against institutional guidelines, not against actual patient outcomes [3]. This highlights a common theme: high accuracy in matching guidelines does not automatically translate to better patient health, especially if the guidelines themselves are imperfect or if the AI's advice is not followed correctly.

Healthcare practitioners themselves have mixed feelings. In interviews with nine NHS emergency department staff about a proposed AI triage system (DAISY), most were positive, citing benefits like reduced wait times and help identifying very sick patients [4]. But they also raised concerns about trust, empathy, and the loss of non-verbal cues in patient interactions [4]. Trust was identified as a 'significant driver of use and a potential barrier to adoption' [4]. This means even a technically excellent AI triage system could fail to improve outcomes if clinicians don't trust or use it properly.

Does AI triage work in places with fewer resources?

The evidence for AI triage in low-resource settings is even thinner. A 2025 article in the World Health Organization's bulletin reviewed the Interagency Integrated Triage Tool, a non-AI triage system used in low- and middle-income countries like Papua New Guinea [6]. The authors found that while the tool can be reliably applied by health workers and reduces waiting times and mortality in pre-post studies, they also noted a 'lack of high-quality evidence supporting an association between triage implementation and improved clinical outcomes' [6]. They cautioned that the existing data 'may have been subject to confounding and publication bias' [6]. This is a sobering reminder: even for established, non-AI triage tools, proving that they actually improve patient outcomes in the real world is difficult. For AI chatbots, which are newer and less tested, the evidence gap is even wider.

About These Sources

This answer is built on 6 peer-reviewed studies — published from 2021 to 2025, 3 from 2024 or later, 3 in Q1 journals, collectively cited 171 times — selected as the most relevant from 6 studies that passed quality screening, drawn from 58 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Real-World Experience with Artificial Intelligence-Based Triage in Transferred Large Vessel Occlusion Stroke Patients

In a retrospective study of 55 stroke patients transferred for endovascular therapy, the Viz LVO AI triage system reduced median neuroendovascular team notification time from 40 to 25 minutes (38% faster) and reduced variation in notification times, though the 25-minute reduction in door-to-puncture time was not statistically significant.

2

Artificial intelligence chatbot performance in triage of ophthalmic conditions

In a cross-sectional study using 44 ophthalmic vignettes, ChatGPT (GPT-4) listed the correct diagnosis in the top three 93% of the time (vs. 95% for physicians) and correctly triaged urgency in 98% of cases (vs. 86% for physicians), with no grossly inaccurate statements.

3

Real-world deployment and evaluation of PEri-operative AI CHatbot (PEACH): a large language model chatbot for peri-operative medicine.

In a silent deployment of the PEACH peri-operative chatbot across 240 real clinical interactions, overall accuracy was 96.7%, improving to 97.9% after updates, with minimal hallucinations (1/240) and deviations (2/240); clinicians reported it expedited decisions in 95% of cases.

4

Medical practitioner perspectives on AI in emergency triage

In qualitative interviews with nine NHS emergency department practitioners about a proposed AI triage system (DAISY), most were positive about potential benefits like reduced wait times and improved identification of very ill patients, but trust was identified as a key barrier to adoption, alongside concerns about empathy and non-verbal cues.

5

AI Triage in Primary Care: Building Safer and More Equitable Real-World Evidence (Preprint)

A 2025 review argues that current evidence for AI triage in primary care is limited to retrospective validations, emergency settings, or vignettes, with almost no real-world outcome data or equity-stratified safety reporting, raising concerns that deployment could widen health inequalities.

6

Triage systems in low-resource emergency care settings

A review of the Interagency Integrated Triage Tool in low-resource settings (e.g., Papua New Guinea) found it can be reliably applied and may reduce waiting times and mortality, but noted a lack of high-quality evidence linking triage implementation to improved clinical outcomes, with potential confounding and publication bias.