Why do AI pathology models perform worse for some patient groups?
The most concerning evidence gap is demographic bias: AI models trained on publicly available datasets can be significantly less accurate for certain racial groups. A 2024 study found that when using common modeling approaches, performance gaps (measured by area under the receiver operating characteristic curve) between white and Black patients were 3.0% for breast cancer subtyping, 10.9% for lung cancer subtyping, and a striking 16.0% for predicting IDH1 mutations in gliomas [2]. These gaps persisted even when using state-of-the-art bias mitigation strategies and richer feature representations from self-supervised models, though the richer models did reduce the disparities [2]. The root cause is that large public datasets like The Cancer Genome Atlas underrepresent certain demographic groups, so models never learn to recognize disease patterns in those populations [2]. This means a model that works well in one hospital system may fail in another with a different patient mix, directly undermining its clinical utility and fairness.
The problem is not limited to race. The same study showed that performance disparities extend to other demographic factors, and the authors explicitly call for regulatory agencies to integrate demographic-stratified evaluation into their assessment guidelines [2]. Without such requirements, developers have little incentive to ensure their models work equally well for everyone.
How much real-world testing have these models actually undergone?
A second critical gap is the near-total absence of prospective clinical trials. A 2021 review noted that 'no randomized prospective trials have yet shown a benefit of AI-based diagnosis' in pathology [4]. This remains largely true as of 2025. A 2024 survey of all 26 AI-based pathology products on the European market found that only 10 (38%) had published peer-reviewed internal validation studies, and only 11 (42%) had peer-reviewed external validation studies [5]. Most products had received regulatory approval via a self-certification route that does not require independent testing [5]. This means that for the majority of commercial AI pathology tools, the evidence supporting their use comes from the developers themselves, not from independent researchers working with real-world clinical data.
The few studies that do exist are mostly retrospective, using archived tissue samples and known outcomes. For example, a large 2023 study developed an AI model to predict treatment response in liver cancer patients using retrospective data from multiple centers, achieving a correlation of r=0.62 between the model's predictions and the actual molecular signature [1]. While promising, this is still a far cry from a prospective trial where the model's predictions would be used to guide real-time treatment decisions. The authors themselves note that external and prospective validation are ongoing [3]. Until such trials are completed and published, clinicians cannot be confident that these models will perform as expected in day-to-day practice.
Can pathologists trust a 'black box' they don't understand?
A third gap is the limited explainability of deep learning models. Many AI pathology systems are 'black boxes' that provide a diagnosis or risk score without showing their reasoning. This is a major barrier to clinical adoption because pathologists need to understand why a model made a particular call in order to trust it and to spot errors. A 2023 review specifically addresses this issue, arguing that explainable AI — methods that make machine learning decisions more transparent — is essential for precision pathology [7]. Without it, pathologists may either blindly accept incorrect AI suggestions or reject useful ones, both of which are dangerous.
The problem is compounded by the fact that AI models can be 'fragile' — they may perform well on the data they were trained on but fail on slightly different images, such as those from a different scanner or staining protocol [8]. A 2025 review notes that challenges include 'limited explainability, data bias, lack of prospective trials, and regulatory hurdles' [6]. Until models can clearly show which features (e.g., specific cell shapes, tissue patterns) drove their decision, pathologists will rightly be cautious about relying on them for patient care.
About These Sources
This answer is built on 8 peer-reviewed studies — published from 2021 to 2025, 4 from 2024 or later, 6 in Q1 journals, collectively cited 480 times — selected as the most relevant from 15 studies that passed quality screening, drawn from 72 papers retrieved from a database of over 500 million.
Sources used in this answer
Artificial intelligence-based pathology as a biomarker of sensitivity to atezolizumab–bevacizumab in patients with hepatocellular carcinoma: a multicentre retrospective study
A 2023 multicenter retrospective study developed an AI model (ABRS-P) that predicted progression-free survival in liver cancer patients treated with atezolizumab-bevacizumab, with a correlation of r=0.62 between model predictions and the molecular signature in the training set, and r=0.53 in an external biopsy validation set.
Demographic bias in misdiagnosis by computational pathology models
A 2024 study found that whole-slide image classification models show marked performance disparities across demographic groups, with gaps between white and Black patients of 3.0% for breast cancer subtyping, 10.9% for lung cancer subtyping, and 16.0% for IDH1 mutation prediction in gliomas.
Development and validation of a multimodal artificial intelligence (MMAI)-derived digital pathology-based biomarker predicting metastasis among patients with biochemical recurrence after radical prostatectomy in NRG/RTOG trials
A 2025 study developed a multimodal AI model predicting distant metastasis in prostate cancer patients after radical prostatectomy, achieving a 10-year time-dependent AUC of 0.74 versus 0.68 for a clinical nomogram, but notes that external and prospective validation are ongoing.
Artificial Intelligence in Pathology
A 2021 review states that no randomized prospective trials have yet shown a benefit of AI-based diagnosis in pathology, and that initial proof-of-concept studies are available but need confirmation or falsification through prospective trials.
Public evidence on AI products for digital pathology
A 2024 review of 26 AI-based pathology products on the European market found that only 10 (38%) had peer-reviewed internal validation studies and only 11 (42%) had peer-reviewed external validation studies; most had regulatory approval via self-certification.
Artificial Intelligence in Clinical Medicine: Challenges Across Diagnostic Imaging, Clinical Decision Support, Surgery, Pathology, and Drug Discovery
A 2025 review of 150 studies across five medical domains concludes that AI shows strong performance in diagnostic imaging (AUC up to 0.94) but faces challenges including limited explainability, data bias, lack of prospective trials, and regulatory hurdles.
Toward Explainable Artificial Intelligence for Precision Pathology
A 2023 review explains the limitations of conventional AI's 'black-box' character and argues that explainable AI is essential for making machine learning decisions more transparent in precision pathology.
AI in Pathology: What could possibly go wrong?
A 2023 review discusses challenges in AI pathology adoption including unrepresentative training data with implicit bias, data privacy concerns, algorithm fragility, and potential for practitioner deskilling and burnout.
