WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

How well does AI triage chatbots handle diverse patient populations?

AI triage chatbots show high accuracy for urgent cases but struggle with diverse populations due to literacy, language, and age barriers.

Direct answer

AI triage chatbots handle diverse patient populations reasonably well for urgent cases, but performance varies significantly by language, literacy, and age. Across the studies here, the larger trials consistently show strong agreement with clinicians—around 80-84% categorical agreement [1][4]—and very high sensitivity for critical patients (92-100%) [3][4]. However, older adults (60+) report much lower satisfaction [1], and systems that don't support local dialects or low-literacy users underperform [4]. The bottom line: these tools are safe and effective as decision-support aids for triaging urgent cases, but they are not yet equally reliable or acceptable across all patient groups.

5sources cited

This article was generated with WisPaper-powered search and paper analysis.

How accurate are AI triage chatbots for different patient groups?

Overall, AI triage chatbots agree with doctors about 80-84% of the time across urgency levels, which is substantial but not perfect. In a real-world UK primary care study of 649 patients, the AI tool (Visiba Triage) showed 83.7% categorical agreement with general practitioners across three urgency levels (kappa 0.69) [1]. A separate study in Morocco using a dialect-native AI chatbot on WhatsApp found 80.7% concordance with physician triage (kappa 0.70) [4]. These two studies, using very different populations and settings, converge on the same level of agreement, which strengthens the finding.

For the most critical patients—those needing immediate emergency care—AI chatbots are highly sensitive, meaning they rarely miss a true emergency. In the Moroccan study, sensitivity for 'urgent' cases was 92.3% [4]. In a study of 46 emergency scenarios, ChatGPT-4o had 100% sensitivity for the most urgent triage level (Level 1) and 100% sensitivity for the least urgent (Level 5), though it was least sensitive (50%) at an intermediate level (Level 4) [3]. This means the tool is very safe for catching the sickest patients, but may sometimes under-triage moderately urgent cases.

However, accuracy drops for certain conditions. A study of AI chatbots for ophthalmic (eye) conditions found that ChatGPT-4 listed the correct diagnosis in its top three suggestions 93% of the time, but Bing Chat managed only 77% and WebMD Symptom Checker just 33% [2]. This wide variation shows that not all AI tools are equal, and performance depends heavily on the specific chatbot and the medical domain.

Who benefits most from AI triage chatbots—and who gets left behind?

The biggest winners are patients who speak a supported language and have basic digital literacy. The Moroccan study deliberately built a system that accepts voice, text, and images in local dialects (Moroccan Darija), English, or French via WhatsApp, and it reduced triage time by nearly 24 minutes on average [4]. This 'zero-friction' design—no app download, no typing required—made it accessible to a broader population in a low-to-middle-income country.

Older adults are the group most likely to be dissatisfied or left behind. In the UK study, patients aged 60 and older were significantly less satisfied with the AI triage tool compared to younger patients (adjusted odds ratio 0.25, meaning they were 75% less likely to report high satisfaction) [1]. This is a critical caveat: even when the AI is clinically accurate, older users may find it confusing or impersonal.

Patients with limited literacy or those who speak unsupported dialects also face barriers. The Moroccan study explicitly notes that traditional AI symptom checkers 'often fail in these settings due to high literacy requirements, lack of dialect support' [4]. So while a dialect-native chatbot can bridge that gap, most existing tools do not offer this feature, leaving large populations underserved.

When should you trust an AI chatbot for triage—and when should you be cautious?

Trust the AI for initial sorting of urgent cases, but always treat it as a decision-support tool, not a replacement for a doctor. Across all studies, the AI was safe for identifying emergencies: in the UK study, no case the AI rated as non-urgent was later reclassified as an emergency by a GP [1]. In the emergency scenarios study, ChatGPT-4o had 100% sensitivity for the most critical patients [3]. This means the AI is very unlikely to send a truly sick person home.

Be cautious with rare or complex conditions. A study on neurogenetic disorders found that ChatGPT and Google's Gemini showed 'notable gaps in diagnostic accuracy and a concerning level of hallucinations'—meaning they sometimes made up plausible-sounding but incorrect information [5]. The authors stress that these tools require 'meticulous collaboration and oversight from both neurologists and geneticists' [5]. For uncommon diseases, AI is far less reliable.

Also be aware that AI tends to over-triage—rating cases as more urgent than they actually are. The UK study noted the AI had a 'safety-conscious design, with a greater likelihood of overtriage' [1], and the ophthalmic study found Bing Chat had 'a tendency to overestimate triage urgency' [2]. While this is safer than under-triaging, it can lead to unnecessary worry and overuse of emergency resources.

About These Sources

This answer is built on 5 peer-reviewed studies — published from 2023 to 2026, 4 from 2024 or later, 1 in Q1–Q2 journals, collectively cited 79 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 50 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Real-World Evaluation of Artificial Intelligence-Assisted Triage for Same-Day Appointments: A Mixed-Methods Study in UK Primary Care (Preprint)

In a real-world UK primary care study of 649 patients, the AI triage tool showed 83.7% categorical agreement with GPs (kappa 0.69), with no missed emergencies, but older adults (60+) were 75% less likely to report high satisfaction.

2

Artificial intelligence chatbot performance in triage of ophthalmic conditions

In a cross-sectional study of 44 ophthalmic vignettes, ChatGPT-4 listed the correct diagnosis in the top three 93% of the time and triaged urgency correctly 98% of the time, outperforming Bing Chat (77% and 84%) and WebMD (33%).

3

Evaluating the Accuracy of Artificial Intelligence Chatbots in Triaging Emergency Cases: A Comparative Study with Expert Clinicians

In a cross-sectional study of 46 emergency scenarios, ChatGPT-4o showed 100% sensitivity for the most urgent triage level (Level 1) and substantial overall agreement with expert physicians (kappa 0.695).

4

Bridging the Literacy Gap in Digital Health: A Prospective Study of a Dialect-Native AI Triage Chatbot for Urgent Care Optimization (Preprint)

In a prospective study of 150 patients in Morocco, a dialect-native AI chatbot on WhatsApp achieved 80.7% concordance with physician triage (kappa 0.70), 92.3% sensitivity for urgent cases, and reduced triage time by 23.8 minutes.

5

AI-Powered Neurogenetics: Supporting Patient's Evaluation with Chatbot.

In a study of 90 questions about six neurogenetic disorders, GPT chatbots outperformed Gemini but all showed notable gaps in diagnostic accuracy and concerning levels of hallucination, requiring clinician oversight.