WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

What evidence gaps are holding back AI triage chatbots?

AI triage chatbots show high accuracy in controlled tests but lack real-world safety and equity data, especially for diverse populations.

Direct answer

AI triage chatbots are surprisingly accurate in controlled tests—one study found ChatGPT matched ophthalmology trainees, listing the correct diagnosis in 93% of cases [1]. But the big gap holding them back is a lack of real-world evidence: almost no studies have tested these chatbots in actual clinics with diverse patients, especially across different ages, ethnicities, and languages [3]. Without this data, we don't know if they're safe and fair for everyone, which is why hospitals are cautious about deploying them. Across the studies here, the strongest evidence comes from controlled vignette studies [1] and a single prospective hospital trial [4], but none have been tested at scale in primary care.

5sources cited

This article was generated with WisPaper-powered search and paper analysis.

What we know: AI chatbots can be surprisingly accurate in controlled tests

In controlled settings, AI triage chatbots perform remarkably well. A 2023 study tested ChatGPT (GPT-4), Bing Chat, and WebMD on 44 ophthalmic case vignettes. ChatGPT listed the correct diagnosis in its top three suggestions 93% of the time, nearly matching ophthalmology trainees (95%), and gave the right triage urgency in 98% of cases—actually outperforming the human doctors on that measure [1]. This shows that, given a clear written description, a top-tier AI can sort patients into the right urgency category with high accuracy.

A 2025 prospective study in Morocco tested a WhatsApp-based AI chatbot that could understand voice messages in local dialects (Moroccan Darija), text, or images. Against physician triage, the AI showed 'substantial agreement' (a statistical measure called Cohen's Kappa of 0.70) and correctly identified 92.3% of urgent cases, with a 97.2% chance that a patient flagged as non-urgent truly didn't need urgent care [4]. This is the first real-world validation of a dialect-native AI triage system, and it cut average triage time by nearly 24 minutes [4].

The critical gap: almost no real-world evidence in primary care or for diverse populations

The biggest evidence gap is simple: almost all studies use made-up cases (vignettes) or emergency department settings, not real patients in primary care clinics. A 2025 review argues that current evidence is 'drawn from retrospective validations, emergency settings, or vignettes, with scant evaluation of real-world outcomes and almost no equity-stratified safety data' [3]. That means we don't know how these chatbots perform when real patients describe symptoms in messy, incomplete ways, or when they have multiple health problems at once.

Even more concerning, there is almost no data on whether AI triage works equally well for different groups of people. The same review points out that known disparities exist 'across age, ethnicity, language, and deprivation,' but AI triage studies have not reported safety data broken down by these factors [3]. Without this, deploying AI triage could accidentally widen health inequalities—for example, by being less accurate for non-native speakers or older adults. A 2025 analysis of AI healthcare startups confirms this, noting that 'real-world evaluations are few' and that adoption is uneven due to 'limited/biased local data' [5].

Trust and adoption are held back by missing evidence on safety and fairness

Medical staff are actually open to AI triage—a 2021 survey of 677 healthcare workers in China found 77% acceptance, with 45% preferring 'AI triage exclusively' [2]. But that acceptance depends on seeing proof that it works. The same survey found that direct experience with AI and exposure to varied media about it increased preference [2]. This suggests that as evidence builds, adoption could grow quickly.

However, the missing evidence creates a chicken-and-egg problem. Without real-world safety and equity data, hospitals can't confidently deploy AI triage at scale. A 2025 analysis warns that risks arise not just from algorithmic bias but also from 'human factors, workflow misalignment, governance gaps, and inadequate postdeployment monitoring' [3]. Until studies track actual patient outcomes—not just accuracy on vignettes—across diverse populations, AI triage chatbots will remain a promising tool that few health systems are willing to bet on.

About These Sources

This answer is built on 5 peer-reviewed studies — published from 2023 to 2025, 3 from 2024 or later, 2 in Q1 journals, collectively cited 92 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 50 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Artificial intelligence chatbot performance in triage of ophthalmic conditions

In a cross-sectional study of 44 ophthalmic vignettes, ChatGPT (GPT-4) listed the correct diagnosis in its top three 93% of the time and gave appropriate triage urgency in 98% of cases, matching or exceeding ophthalmology trainees (95% and 86%, respectively).

2

AI triage or manual triage? Exploring medical staffs’ preference for AI triage in China

A survey of 677 medical staff in China found 77.1% acceptance of AI triage, with 45.2% preferring exclusive AI triage; direct experience and media exposure were positively associated with preference.

3

AI Triage in Primary Care: Building Safer and More Equitable Real-World Evidence (Preprint)

A 2025 review argues that current AI triage evidence is limited to retrospective validations, emergency settings, or vignettes, with no real-world primary care evaluations and almost no equity-stratified safety data across age, ethnicity, language, or deprivation.

4

Bridging the Literacy Gap in Digital Health: A Prospective Study of a Dialect-Native AI Triage Chatbot for Urgent Care Optimization (Preprint)

In a prospective study of 150 patients at a Moroccan hospital, a WhatsApp-based AI triage chatbot (using GPT-4 and Whisper) showed substantial agreement with physician triage (Kappa=0.70), 92.3% sensitivity for urgent cases, and reduced mean triage time by 23.8 minutes.

5

AI-Driven Healthcare Entrepreneurship: Transforming Clinical Practice Through Innovation, Access, and Affordability

A 2025 mixed-methods synthesis of AI healthcare startups found that real-world evaluations of AI triage chatbots are few, cost-effectiveness is context-dependent, and adoption is hindered by infrastructure gaps, biased local data, and trust barriers.