How much workload can AI actually cut?
The most reliable evidence from these studies points to a workload reduction of roughly 15–18% when AI is used as a triage or assistance tool. In a multicenter prospective trial of over 1,000 chest X-rays, AI assistance reduced radiologist interpretation time by 18.5% (from an average of 4.11 to 4.37 quality score, P < 0.001) [5]. Another study of a commercial AI system on 1,670 chest radiographs found that by safely excluding half of all normal exams from reporting—with a 98% negative predictive value for urgent findings—workload dropped by 15% [2]. These figures come from different designs (prospective trial vs. retrospective analysis) but converge on the same range, which strengthens the conclusion.
The workload savings come from letting AI handle the easy cases first. In the triage approach, the AI identifies normal or near-normal studies and either auto-reports them or removes them from the reading queue, so radiologists only see the abnormal or uncertain ones. This is fundamentally different from having AI generate a full report that a radiologist must then check—which can sometimes add time rather than save it.
Does accuracy suffer when AI helps?
No—in fact, accuracy often improves when AI is used as a copilot rather than a replacement. The same prospective trial that found an 18.5% time reduction also found that AI-assisted reports had significantly higher quality scores (4.37 vs. 4.11 out of 5, P < 0.001) [5]. A separate study using a multi-model fusion framework (combining ChatGPT and Claude) achieved 91.3% accuracy on chest X-ray interpretation when the two models agreed, compared to 84% for the best single model [4]. This consensus approach reduced diagnostic errors without adding clinician time.
However, the evidence is clear that AI alone is not reliable enough to work unsupervised. One study tested a single-prompt LLM for detecting errors in radiology reports and found a positive predictive value (PPV) of only 6.3%—meaning 94% of its flagged errors were false alarms [1]. A three-pass framework improved that to 15.9% PPV, but that still means most flags are wrong. This is why the safe deployment model is AI-as-triage or AI-as-second-reader, not AI-as-final-reader.
A 2022 review of the broader literature on radiologist workload and error rates concluded that the scientific evidence needed to set safe workload limits is still lacking, and that regulating workloads without evidence could be more harmful than not regulating at all [6]. This underscores that AI copilots are a promising tool, but they are not a substitute for understanding the fundamental relationship between speed, volume, and accuracy in radiology.
Who benefits most from AI copilots?
The benefits are largest in high-volume, resource-constrained settings—especially primary care and emergency departments where chest X-rays are the most common exam. The multicenter trial specifically highlighted that its lightweight AI system (1 billion parameters) was designed for resource-constrained settings and outperformed much larger models like ChatGPT 4o (200 billion parameters) on chest X-ray report generation [3][5]. This means even clinics without expensive hardware can deploy effective AI assistance.
Radiologists reading high volumes of normal exams benefit most. The triage study showed that 29% of all chest radiographs in the dataset were normal (479 out of 1,670), and the AI could safely remove half of those from the reading queue [2]. For a radiologist seeing 100 chest X-rays per shift, that means 15 fewer studies to read—time that can be spent on complex cases. The error-detection framework also showed that AI could reduce the number of reports needing human review from 192 to 88 per 1,000 reports, cutting operational costs by 42.6% [1].
On the other hand, radiologists working in highly specialized settings (e.g., neuroimaging or oncology follow-up) may see less benefit, because the AI systems studied here are optimized for chest X-rays and general radiology reports. The evidence does not yet extend to all imaging modalities or all clinical contexts.
About These Sources
This answer is built on 6 studies (4 peer-reviewed, 2 preprints) — published from 2022 to 2025, 5 from 2024 or later, 3 in Q1 journals, collectively cited 122 times — selected as the most relevant from 7 studies that passed quality screening, drawn from 50 papers retrieved from a database of over 500 million.
Sources used in this answer
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
A three-pass LLM framework for detecting errors in radiology reports improved positive predictive value from 6.3% to 15.9% and reduced operational costs by 42.6% compared to a single-prompt approach, while maintaining stable detection rates (absolute true positive rate 0.012–0.014) in a retrospective analysis of 1,000 reports from MIMIC-III.
Performance of AI to exclude normal chest radiographs to reduce radiologists’ workload
A commercial AI system identified normal chest radiographs with an AUC of 0.92 and, at a conservative operating point, could exclude 53% of normal exams from reporting with a 98% negative predictive value for urgent findings, yielding a 15% workload reduction in a retrospective analysis of 1,670 chest radiographs.
A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice
In a multicenter prospective trial, the Janus-Pro-CXR system (1B parameters) reduced radiologist interpretation time by 18.3% (P < 0.001) while improving report quality scores, and outperformed ChatGPT 4o (200B parameters) on automated chest X-ray report generation.
Fusion-Augmented Large Language Models: Boosting Diagnostic Trustworthiness via Model Consensus
A multi-model fusion framework using ChatGPT and Claude achieved 91.3% accuracy on chest X-ray interpretation when the two models agreed (95% similarity threshold), compared to 84% for the best single model, in a subset of 50 multimodal cases from the CheXpert dataset.
From Bench to Bedside: A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice
In a multicenter prospective trial, the DeepSeek-powered Janus-Pro-CXR system reduced interpretation time by 18.5% (P < 0.001), improved report quality scores (4.37 vs. 4.11, P < 0.001), and was preferred by a majority of experts in 52.7% of cases, while detecting eight critical findings with AUC > 0.8.
Mandating Limits on Workload, Duty, and Speed in Radiology
A 2022 review concluded that the scientific evidence needed to set safe workload or duty limits for radiologists is lacking, and that regulating workloads without evidence could be more harmful than not regulating at all.
