How does AI compare to radiologists working alone?
The largest and most comprehensive study here, involving 20 radiologists and over 2,500 chest X-rays, found that the AI model alone was significantly more accurate than unassisted radiologists for 117 out of 124 clinical findings (94%), and was non-inferior for the remaining findings [2]. The AI's macro-averaged area under the curve (AUC) was 0.957, compared to 0.713 for unassisted radiologists — a huge difference [2]. However, an international reader study with 204 radiologists from 11 countries found the opposite: radiologists had a higher overall AUC (0.96 vs. 0.87) and better sensitivity (0.91 vs. 0.71) than AI [1]. The key difference is that the first study tested AI on a very broad set of 127 findings, while the second used a more routine, general dataset. So AI can outperform radiologists on specific, well-defined tasks, but radiologists still hold the edge on general, unselected cases.
Can AI help radiologists do better?
Yes, and the evidence is strong and consistent. In the same large study, when radiologists used the AI as a decision-support tool, their accuracy improved significantly for 102 out of 127 findings (80%), and no finding showed a decrease in accuracy [2]. Another study with 8 radiologists found that referring to a deep learning model significantly improved their ability to distinguish COVID-19 from other pneumonias (p = 0.0038) [8]. A separate study showed that both radiologists and non-radiologist physicians significantly improved their chest X-ray abnormality detection when aided by an FDA-cleared AI system, with non-radiologists matching radiologist accuracy when using the AI [4]. The AI also caught errors: in a multicenter study of 279 cases with missed or mislabeled findings, the AI detected them with 96% sensitivity and 100% specificity [5]. Across these studies, AI consistently acts as a safety net, catching what humans miss and boosting overall accuracy.
Where does AI still fall short?
AI is not perfect, and the evidence highlights several limitations. First, AI struggles with small or complex pathologies: a benchmarking study of saliency methods (which show where the AI is 'looking') found that all seven methods performed significantly worse than human experts at localizing pathologies, especially for smaller and more irregularly shaped findings [7]. Second, AI can be biased: one study found that diagnostic AI in a large public dataset had a higher false negative rate for racial minorities, though oversampling and data augmentation reduced this disparity by up to 74.7% without hurting overall performance [9]. Third, in a hospital-at-home setting, agreement between AI and radiologists was only moderate (kappa = 0.49), and the authors concluded AI needs further development before clinical use [6]. Finally, while AI matches or beats less experienced radiologists, it still lags behind experienced specialists in overall accuracy [1][3]. The takeaway: AI is a powerful tool, but it is not yet a standalone replacement for expert human judgment.
About These Sources
This answer is built on 9 peer-reviewed studies — published from 2021 to 2025, 4 from 2024 or later, 6 in Q1 journals, collectively cited 449 times — selected as the most relevant from 11 studies that passed quality screening, drawn from 47 papers retrieved from a database of over 500 million.
Sources used in this answer
An International Non-Inferiority Study for the Benchmarking of AI for Routine Radiology Cases: Chest X-ray, Fluorography and Mammography
In an international reader study with 204 radiologists, AI had a lower AUROC (0.87 vs. 0.96) and lower sensitivity (0.71 vs. 0.91) than radiologists, but was non-inferior to the least experienced radiologists for chest X-ray, suggesting AI could be used for first reading to reduce workload.
Effect of a comprehensive deep-learning model on the accuracy of chest x-ray interpretation by radiologists: a retrospective, multireader multicase study
In a large multireader study (20 radiologists, 2,568 cases), AI alone was significantly more accurate than unassisted radiologists for 94% of 124 findings (AUC 0.957 vs. 0.713), and when used as decision support, it improved radiologist accuracy for 80% of findings.
Comparing the accuracy of computer-aided detection (CAD) software and radiologists from multiple countries for tuberculosis detection in chest X-Rays.
Comparing 12 CAD software to 11 radiologists from four countries on 774 chest X-rays, the top CAD outperformed all radiologists except Indian radiologists when matching specificity, and British radiologists' sensitivity was closest to the top CAD.
Deep learning improves physician accuracy in the comprehensive detection of abnormalities on chest X-rays
An FDA-cleared AI system (AUC 0.976) significantly improved physician accuracy (difference in AUC 0.101, p<0.001), and non-radiologist physicians using AI matched radiologist accuracy and were faster.
Performance of a Chest Radiography AI Algorithm for Detection of Missed or Mislabeled Findings: A Multicenter Study
In a multicenter study of 279 chest X-rays with missed or mislabeled findings, an AI algorithm detected these errors with 96% sensitivity, 100% specificity, and 96% accuracy, showing potential to reduce radiologist errors.
Consensus Between Radiologists, Specialists in Internal Medicine, and AI Software on Chest X-Rays in a Hospital-at-Home Service: Prospective Observational Study.
In a hospital-at-home service, agreement between AI and radiologists was moderate (kappa=0.49), while agreement between internal medicine specialists and radiologists was substantial (kappa=0.65), suggesting AI needs further development for clinical use.
Benchmarking saliency methods for chest X-ray interpretation
Benchmarking seven saliency methods (including Grad-CAM) against a human expert benchmark found all methods performed significantly worse at localizing pathologies, especially for smaller and more complex shapes, and model confidence was positively correlated with localization performance.
Computer-aided diagnosis of chest X-ray for COVID-19 diagnosis in external validation study by radiologists with and without deep learning system
A deep learning model for COVID-19 had better diagnostic performance than most radiologists (accuracy 0.733 vs. 0.696), and significantly improved radiologist performance when used as a reference (p=0.0038).
Enhancement of Fairness in AI for Chest X-ray Classification.
In the MIMIC-CXR dataset, AI had a higher false negative rate for racial minorities; oversampling reduced this disparity by 74.7% without reducing performance (AUC 0.816-0.820 vs. baseline 0.810-0.819).
