What we know: AI chatbots can be surprisingly accurate in controlled tests
In controlled settings, AI triage chatbots perform remarkably well. A 2023 study tested ChatGPT (GPT-4), Bing Chat, and WebMD on 44 ophthalmic case vignettes. ChatGPT listed the correct diagnosis in its top three suggestions 93% of the time, nearly matching ophthalmology trainees (95%), and gave the right triage urgency in 98% of cases—actually outperforming the human doctors on that measure [1]. This shows that, given a clear written description, a top-tier AI can sort patients into the right urgency category with high accuracy.
A 2025 prospective study in Morocco tested a WhatsApp-based AI chatbot that could understand voice messages in local dialects (Moroccan Darija), text, or images. Against physician triage, the AI showed 'substantial agreement' (a statistical measure called Cohen's Kappa of 0.70) and correctly identified 92.3% of urgent cases, with a 97.2% chance that a patient flagged as non-urgent truly didn't need urgent care [4]. This is the first real-world validation of a dialect-native AI triage system, and it cut average triage time by nearly 24 minutes [4].
The critical gap: almost no real-world evidence in primary care or for diverse populations
The biggest evidence gap is simple: almost all studies use made-up cases (vignettes) or emergency department settings, not real patients in primary care clinics. A 2025 review argues that current evidence is 'drawn from retrospective validations, emergency settings, or vignettes, with scant evaluation of real-world outcomes and almost no equity-stratified safety data' [3]. That means we don't know how these chatbots perform when real patients describe symptoms in messy, incomplete ways, or when they have multiple health problems at once.
Even more concerning, there is almost no data on whether AI triage works equally well for different groups of people. The same review points out that known disparities exist 'across age, ethnicity, language, and deprivation,' but AI triage studies have not reported safety data broken down by these factors [3]. Without this, deploying AI triage could accidentally widen health inequalities—for example, by being less accurate for non-native speakers or older adults. A 2025 analysis of AI healthcare startups confirms this, noting that 'real-world evaluations are few' and that adoption is uneven due to 'limited/biased local data' [5].
Trust and adoption are held back by missing evidence on safety and fairness
Medical staff are actually open to AI triage—a 2021 survey of 677 healthcare workers in China found 77% acceptance, with 45% preferring 'AI triage exclusively' [2]. But that acceptance depends on seeing proof that it works. The same survey found that direct experience with AI and exposure to varied media about it increased preference [2]. This suggests that as evidence builds, adoption could grow quickly.
However, the missing evidence creates a chicken-and-egg problem. Without real-world safety and equity data, hospitals can't confidently deploy AI triage at scale. A 2025 analysis warns that risks arise not just from algorithmic bias but also from 'human factors, workflow misalignment, governance gaps, and inadequate postdeployment monitoring' [3]. Until studies track actual patient outcomes—not just accuracy on vignettes—across diverse populations, AI triage chatbots will remain a promising tool that few health systems are willing to bet on.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2023 to 2025, 3 from 2024 or later, 2 in Q1 journals, collectively cited 92 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 50 papers retrieved from a database of over 500 million.
Sources used in this answer
Artificial intelligence chatbot performance in triage of ophthalmic conditions
In a cross-sectional study of 44 ophthalmic vignettes, ChatGPT (GPT-4) listed the correct diagnosis in its top three 93% of the time and gave appropriate triage urgency in 98% of cases, matching or exceeding ophthalmology trainees (95% and 86%, respectively).
AI triage or manual triage? Exploring medical staffs’ preference for AI triage in China
A survey of 677 medical staff in China found 77.1% acceptance of AI triage, with 45.2% preferring exclusive AI triage; direct experience and media exposure were positively associated with preference.
AI Triage in Primary Care: Building Safer and More Equitable Real-World Evidence (Preprint)
A 2025 review argues that current AI triage evidence is limited to retrospective validations, emergency settings, or vignettes, with no real-world primary care evaluations and almost no equity-stratified safety data across age, ethnicity, language, or deprivation.
Bridging the Literacy Gap in Digital Health: A Prospective Study of a Dialect-Native AI Triage Chatbot for Urgent Care Optimization (Preprint)
In a prospective study of 150 patients at a Moroccan hospital, a WhatsApp-based AI triage chatbot (using GPT-4 and Whisper) showed substantial agreement with physician triage (Kappa=0.70), 92.3% sensitivity for urgent cases, and reduced mean triage time by 23.8 minutes.
AI-Driven Healthcare Entrepreneurship: Transforming Clinical Practice Through Innovation, Access, and Affordability
A 2025 mixed-methods synthesis of AI healthcare startups found that real-world evaluations of AI triage chatbots are few, cost-effectiveness is context-dependent, and adoption is hindered by infrastructure gaps, biased local data, and trust barriers.
