How often do AI triage chatbots get it wrong, and what kind of mistakes do they make?
The biggest safety risk isn't that chatbots miss emergencies—it's that they cry wolf too often. In a 2025 study of 138 emergency department patients, ChatGPT agreed with the final consultant decision only 42.86% of the time, and it consistently assigned higher urgency levels than the consultant did [1]. That means for every 10 patients, the chatbot would have sent about 6 to a higher level of care than needed. Over-triaging like this can overwhelm emergency rooms, delay care for truly sick patients, and cause unnecessary stress and expense for patients.
A separate 2023 study on eye conditions found that ChatGPT (using GPT-4) got triage urgency right 98% of the time—matching human doctors—but another chatbot, Bing Chat, was correct only 84% of the time and also tended to overestimate urgency [2]. The key takeaway: performance varies wildly by the specific AI model, and even the best ones can't be trusted blindly.
Can AI chatbots handle complex or subtle medical cases safely?
Not reliably, according to the evidence. A 2026 systematic review of eight studies on AI chatbots for primary care triage concluded that chatbots have "limitations in complex reasoning, inconsistent handling of nuanced presentations, and lack of access to non-verbal cues"—all of which are central to safe triage [3]. In other words, a chatbot can't see that a patient looks pale, is breathing fast, or is in obvious distress, which are often the real clues to a serious problem.
The same review found that chatbots sometimes produce "hallucinations"—confidently stated but completely wrong information—and that their accuracy varies by the clinical complexity of the case [3]. The emergency department study also noted that ChatGPT performed better on mid-range acuity cases but was less reliable at the extremes (very mild or very severe) [1]. So for straightforward, textbook symptoms, a chatbot might be helpful; for anything unusual or multi-layered, it's a gamble.
What does this mean for patients and clinicians right now?
The bottom line is that AI triage chatbots are best used as a support tool, not a replacement for a human expert. All three studies agree on this point. The emergency department study explicitly says AI "could support ED triage" but warns that its tendency to overestimate severity could lead to "over-triaging and increased resource use" [1]. The eye-condition study found that ChatGPT performed well but still recommends human oversight [2]. And the systematic review concludes that chatbot triage "is best deployed as clinician-supervised decision support rather than a replacement for professional assessment" [3].
For patients, this means you should not rely on a chatbot to decide whether you need emergency care. If you're worried, see a doctor. For hospitals and clinics, the evidence suggests that deploying an AI triage tool without a human in the loop—at least for now—would be unsafe. The risks of over-triaging, missing subtle cues, and generating incorrect information are real and not yet solved.
About These Sources
This answer is built on 3 peer-reviewed studies — published from 2023 to 2026, 2 from 2024 or later, 1 in Q1–Q2 journals, collectively cited 78 times — selected as the most relevant from 3 studies that passed quality screening, drawn from 46 papers retrieved from a database of over 500 million.
Sources used in this answer
Safety and accuracy of AI in triaging patients in the emergency department.
In a prospective study of 138 emergency department patients, ChatGPT agreed with physicians 85.6% of the time but with the final consultant only 42.86%, consistently overestimating severity and risking overtriage.
Artificial intelligence chatbot performance in triage of ophthalmic conditions
In a cross-sectional study of 44 ophthalmic vignettes, ChatGPT (GPT-4) matched physician triage urgency 98% of the time, while Bing Chat was correct 84% and made some grossly inaccurate statements.
AI Chatbots for Primary Care Triage: A systematic review
A systematic review of 8 studies found that AI chatbots show variable accuracy, struggle with complex reasoning and non-verbal cues, and risk hallucinations, supporting their use only as supervised decision support.
