How accurate are AI pathology models—and does accuracy equal safety?
The models are remarkably accurate, but accuracy alone doesn't guarantee safety. A 2021 Nature study trained a deep-learning model (TOAD) to identify the origin of cancers from routine pathology slides: it achieved 83% top-1 accuracy and 96% top-3 accuracy on known tumors, and 80% and 93% on an external test set [3]. That means in 4 out of 5 cases, the model's single best guess matched the actual cancer origin—a level of performance that could genuinely assist pathologists. However, the same study notes that for 317 cases of cancer of unknown primary, the model's top prediction agreed with the final diagnosis only 61% of the time [3]. So while the model is powerful, it still makes mistakes in a significant minority of cases, and those mistakes could lead to wrong treatment decisions if the model is used without human oversight.
A more recent 2024 model, CONCH, goes even further—it's a visual-language foundation model trained on over 1.17 million image-caption pairs and achieved state-of-the-art results on 14 different pathology tasks, including image classification, segmentation, and captioning [1]. This breadth means a single model could be used for many different diagnostic tasks, which increases both its utility and its risk: a flaw in the model could propagate across multiple clinical decisions. The key point is that these models are not infallible, and their high accuracy can create a false sense of security.
What are the specific safety risks that might be underestimated?
The biggest risks are not technical failures but ethical and systemic ones. A 2021 review in The American Journal of Pathology explicitly warns that AI, if used unethically, 'may exacerbate existing inequities of health care' [2]. The review identifies three foundational principles for safe AI in pathology: transparency (so clinicians understand how a model reaches a decision), accountability (clear lines of responsibility when a model makes an error), and governance (rules and oversight for how AI is deployed) [2]. Without these, even an accurate model can cause harm—for example, by being trained on data that doesn't represent all patient populations, leading to worse performance for certain groups.
A 2025 paper on AI in toxicologic pathology (used in drug safety testing) adds that AI applications are now being classified by their risk level, with 'nonexempt' applications requiring special scrutiny [4]. This suggests that regulators are beginning to recognize that some AI uses carry higher safety stakes. The risk is especially high in pathology because decisions directly affect diagnosis and treatment—a wrong AI output could lead to a missed cancer or an incorrect therapy. The 2021 renal pathology review also stresses that close collaboration between computer scientists and pathologists is essential to ensure AI innovations are actually translatable to clinical practice [5]. If these collaborations are weak, safety gaps widen.
Who benefits most from AI pathology models, and under what conditions are they safe?
The patients who stand to benefit most are those with hard-to-diagnose cancers, such as cancer of unknown primary (CUP). The TOAD model was specifically designed for these cases and achieved top-3 agreement with the final diagnosis in 82% of CUP cases [3]. That means the model can narrow down the possible cancer origin to three options in over 4 out of 5 cases, which can guide further testing and treatment. However, the model is safest when used as an 'assistive tool' alongside a pathologist, not as a standalone decision-maker [3]. The same principle applies to the CONCH model: its ability to handle both images and text makes it versatile, but the authors note it can be used for workflows requiring 'minimal or no further supervised fine-tuning' [1]—which is efficient but also means less human oversight if deployed carelessly.
The conditions for safe use are clear from the evidence: models must be transparent about their limitations, pathologists must remain in the loop, and deployment must be governed by ethical principles [2]. The 2021 renal pathology review adds that AI applications must be 'renal pathology-optimized'—meaning they need to be tailored to the specific medical context, not just generic AI tools [5]. In short, the benefits are real and substantial, but they depend on careful implementation, not just on the model's technical performance.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2021 to 2025, 2 from 2024 or later, 4 in Q1 journals, collectively cited 1,274 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 45 papers retrieved from a database of over 500 million.
Sources used in this answer
A visual-language foundation model for computational pathology
CONCH, a visual-language foundation model trained on over 1.17 million image-caption pairs, achieved state-of-the-art performance across 14 diverse pathology benchmarks, including image classification, segmentation, captioning, and retrieval tasks [1].
Ethics of AI in Pathology
This review identifies transparency, accountability, and governance as three foundational principles for ethical AI in pathology, warning that without them AI could exacerbate healthcare inequities [2].
AI-based pathology predicts origins for cancers of unknown primary
The TOAD deep-learning model achieved 83% top-1 and 96% top-3 accuracy in identifying cancer origins from histology slides, and 61% top-1 concordance for cancers of unknown primary [3].
Digital Pathology and Artificial Intelligence Applied to Nonclinical Toxicology Pathology—The Current State, Challenges, and Future Directions
This 2025 paper on AI in toxicologic pathology notes that nonexempt AI applications are now classified by risk level, indicating growing regulatory attention to safety [4].
AI applications in renal pathology
This review emphasizes that successful AI in renal pathology requires close interdisciplinary collaboration between computer scientists and pathologists to ensure clinical translatability [5].
