WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

Are the safety risks of medical foundation models being underestimated?

Evidence shows medical foundation models hallucinate at alarming rates, with specialized models often worse than general ones, posing real patient harm risks.

Direct answer

Yes, the safety risks of medical foundation models are likely being underestimated. A 2025 study of 11 models found that even the best general-purpose models hallucinated (produced factually incorrect or clinically misleading outputs) in about 12% of cases, while specialized medical models failed in up to 72% of responses [1]. A survey of 70 clinicians in that same study revealed that 91.8% had encountered such hallucinations, and 84.7% believed they could cause patient harm [1]. The problem is not just about knowledge gaps—most dangerous errors stem from faulty reasoning, which current safety testing may not catch.

5sources cited

This article was generated with WisPaper-powered search and paper analysis.

How often do these models get things wrong, and does specialization help?

A 2025 study systematically tested 11 foundation models—7 general-purpose (like GPT-4, Gemini) and 4 built specifically for medicine (like MedGemma)—across seven medical reasoning and information retrieval tasks [1]. The results were striking: general-purpose models produced hallucination-free responses a median of 76.6% of the time, while medical-specialized models managed only 51.3%—a 25 percentage point gap [1]. This means specialized medical models were actually worse, not better, at avoiding dangerous errors. The top performer, Gemini-2.5 Pro, reached 87.6% accuracy on its own and over 97% when using chain-of-thought reasoning (a step-by-step thinking process) [1]. In contrast, the medical model MedGemma scored as low as 28.6% on some tasks [1]. The takeaway: training on medical data alone does not guarantee safety; sophisticated reasoning ability matters more.

What kinds of errors are most dangerous, and how often do they happen?

The same 2025 study had physicians audit the remaining hallucinations after best-practice prompting [1]. They found that 64–72% of residual errors were not due to missing medical knowledge but to failures in causal or temporal reasoning—for example, confusing the order of symptoms and treatments or mistaking correlation for causation [1]. This is critical because such errors can lead to wrong diagnoses or treatment plans that seem plausible but are clinically backwards. The survey of 70 clinicians confirmed the real-world impact: 91.8% had seen medical hallucinations in practice, and 84.7% judged them capable of causing patient harm [1]. These numbers suggest that current safety testing, which often focuses on factual accuracy, may miss the most insidious failure modes.

Do image-based models (like for pathology) have the same risks?

A 2023 study developed a pathology image model trained on 208,414 images from medical Twitter and achieved state-of-the-art performance on classifying new images, with F1 scores (a measure of accuracy) ranging from 0.565 to 0.832 in zero-shot tasks—far better than the 0.030–0.481 of previous models [2]. However, even the best score of 0.832 means roughly 17% of classifications were incorrect. The study did not directly measure hallucinations or clinical reasoning errors, but the risk remains: a model that misclassifies a pathology slide could lead to a missed cancer diagnosis or unnecessary treatment [2]. A 2023 perspective article in Nature warned that as these models become more general and flexible—interpreting images, lab results, and text together—they will challenge current regulatory and validation frameworks, which are not designed for such complex, multi-modal outputs [4]. The European AI Act, as analyzed in a 2023 legal paper, may also leave gaps: even low-risk AI systems can still cause harm to fundamental rights and user safety, and the Act's coverage may be broader in text than in practical effect [5].

About These Sources

This answer is built on 5 peer-reviewed studies — published from 2023 to 2025, 1 from 2024 or later, 3 in Q1 journals, collectively cited 2,111 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 67 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Medical Hallucination in Foundation Models and Their Impact on Healthcare

In a 2025 evaluation of 11 foundation models across seven medical tasks, general-purpose models outperformed medical-specialized models (median hallucination-free rate 76.6% vs 51.3%), and physician audits showed 64–72% of residual errors stemmed from reasoning failures, not knowledge gaps; a survey of 70 clinicians found 91.8% had encountered hallucinations and 84.7% believed they could cause patient harm.

2

A visual–language foundation model for pathology image analysis using medical Twitter

A 2023 study trained a pathology image model (PLIP) on 208,414 images from medical Twitter, achieving state-of-the-art zero-shot F1 scores of 0.565–0.832 across four external datasets, but even the best score implies roughly 17% misclassification risk.

3

On the challenges and perspectives of foundation models for medical image analysis

A 2023 review article discusses the potential of medical foundation models for image analysis, noting they can reduce dependence on labeled data and improve diagnosis, but also highlights challenges including safety and the need for robust validation.

4

Foundation models for generalist medical artificial intelligence

A 2023 perspective in Nature proposes generalist medical AI (GMAI) that will flexibly interpret multiple data types and produce expressive outputs, but warns that such models will challenge current regulatory and validation strategies for medical AI devices.

5

A Blanket That Leaves the Feet Cold: Exploring the AI Act Safety Framework for Medical AI

A 2023 legal analysis of the EU AI Act argues that even minimal-risk AI systems can still cause harm to fundamental rights and user safety, and that the Act's practical coverage may be more limited than its textual scope suggests.