WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

Can AI triage chatbots reduce clinician workload without increasing errors?

AI triage chatbots can cut clinician workload by up to 70% in screening, but accuracy drops in complex diagnostic uncertainty cases.

Direct answer

Yes, AI triage chatbots can reduce clinician workload without increasing errors, but only in well-defined, high-volume screening tasks. In breast cancer screening, an AI triage strategy cut radiologist workload by 72.5% while maintaining noninferior cancer detection [1]. However, in emergency triage, ChatGPT-4o showed only moderate agreement with expert physicians (kappa 0.695) and performed worse than residents on questions involving diagnostic uncertainty [2][4]. The evidence across these studies consistently shows that AI triage works best when the task is narrow and the data is structured, but it still falls short of human judgment in complex or ambiguous cases.

5sources cited

This article was generated with WisPaper-powered search and paper analysis.

The best case: AI triage can slash workload without harming outcomes

In high-volume, standardized screening programs, AI triage can dramatically reduce the number of cases that need human review while keeping detection rates steady. A 2021 study of nearly 16,000 breast cancer screening exams found that an AI triage system could cut radiologist reading time by 72.5% compared to double reading of digital breast tomosynthesis (DBT) images — from 568 hours down to 156 hours [1]. At the same time, the AI detected 95 of 113 cancers versus 92 detected by double reading, a difference that was not statistically significant, meaning cancer detection was noninferior [1]. The recall rate (false alarms) also dropped by 16.7% [1]. This is the strongest quantitative evidence among these studies that AI can meaningfully reduce workload without increasing errors — but it applies to a very specific task: triaging mammograms into low-risk (skip human review) and high-risk (send to radiologist) categories.

The catch: AI struggles with diagnostic uncertainty and complex cases

When the clinical picture is ambiguous — the hallmark of many real-world triage decisions — AI chatbots underperform compared to trained clinicians. A 2024 study tested two leading chatbots (GPT-4o and Claude-3) on questions designed to involve diagnostic uncertainty, where symptoms and history don't point to a clear diagnosis. Family medicine residents scored 61-63% correct, while Claude-3 scored 57.7% and GPT-4o scored 53.3% — both significantly worse [2]. Most of GPT-4o's errors were logical errors (62.5%), meaning it made reasoning mistakes rather than missing facts [2]. Similarly, in emergency triage, ChatGPT-4o showed only moderate agreement with expert physicians (kappa 0.695, where 1.0 is perfect agreement) and was least sensitive (50%) for mid-urgency cases (Triage Level 4), meaning it often under-triaged patients who actually needed attention [4]. These results show that AI triage is not yet reliable enough to replace human judgment in complex or borderline cases.

The gap between potential and reality

The studies paint a clear picture: AI triage can excel at narrow, repetitive tasks like reading mammograms or sorting obvious emergencies, but it falters when nuance is required. In ophthalmology triage, ChatGPT-4 listed the correct diagnosis in its top three 93% of the time and gave appropriate triage urgency 98% of the time — comparable to physicians [3]. Yet in the same study, Bing Chat was less accurate (77% correct diagnosis) and tended to overestimate urgency, showing that not all AI chatbots are equal [3]. A 2026 paper on an AI-powered e-triage system deployed in UAE hospitals reported a 45% reduction in documentation time and 82% improvement in prioritization accuracy, but it also acknowledged ongoing challenges with data privacy, algorithmic bias, and integration with existing electronic health records [5]. The takeaway: AI triage can reduce workload in specific, well-defined settings, but deploying it broadly without human oversight risks missing critical cases or misclassifying patients.

About These Sources

This answer is built on 5 peer-reviewed studies — published from 2021 to 2026, 3 from 2024 or later, 1 in Q1 journals, collectively cited 247 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 50 papers retrieved from a database of over 500 million.

Sources used in this answer

1

AI-based Strategies to Reduce Workload in Breast Cancer Screening with Mammography and Tomosynthesis: A Retrospective Evaluation

In a retrospective study of 15,987 breast cancer screening exams, an AI triage strategy reduced radiologist workload by 72.5% while maintaining noninferior cancer detection (95 vs 92 cancers detected) and reducing recall rate by 16.7% [1].

2

The future of AI clinicians: assessing the modern standard of chatbots and their approach to diagnostic uncertainty.

On diagnostic uncertainty questions, GPT-4o (53.3% correct) and Claude-3 (57.7%) both scored significantly lower than family medicine residents (61-63%), with most of GPT-4o's errors being logical reasoning mistakes [2].

3

Artificial intelligence chatbot performance in triage of ophthalmic conditions

In a cross-sectional study of 44 ophthalmic vignettes, ChatGPT-4 listed the correct diagnosis in its top three 93% of the time and gave appropriate triage urgency 98% of the time, comparable to physicians, while Bing Chat was less accurate (77%) and overestimated urgency [3].

4

Evaluating the Accuracy of Artificial Intelligence Chatbots in Triaging Emergency Cases: A Comparative Study with Expert Clinicians.

In a study of 46 emergency cases, ChatGPT-4o showed moderate agreement with expert physicians (kappa 0.695), with 100% sensitivity for the most critical cases but only 50% sensitivity for mid-urgency cases (Triage Level 4) [4].

5

AI-Powered E-Triage Systems: Enhancing Healthcare Efficiency Through Intelligent Clinical Decision Support

An AI-powered e-triage system deployed in UAE hospitals reduced documentation time by 45% and improved patient prioritization accuracy by 82%, but the paper notes ongoing challenges with data privacy, algorithmic bias, and EHR integration [5].