WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

Do medical foundation models have enough prospective clinical evidence?

Prospective clinical evidence for medical foundation models is rare; only one RCT exists, showing clear benefit for an eyecare copilot.

Direct answer

No, most medical foundation models still lack robust prospective clinical evidence. Across the studies reviewed here, only one randomized controlled trial (RCT) has been completed — an eyecare copilot called EyeFM — which showed that ophthalmologists using it achieved a 92.2% correct diagnostic rate versus 75.4% without it [1]. The vast majority of other models have only been validated retrospectively or in small, single-site studies, and fewer than 5% of pathology foundation model studies have conducted prospective trials [7]. So while the technical performance is often excellent in controlled settings, the evidence needed to prove they improve real-world patient outcomes is still very thin.

7sources cited

This article was generated with WisPaper-powered search and paper analysis.

What is the single best piece of prospective evidence?

The strongest prospective clinical evidence for any medical foundation model comes from a single randomized controlled trial (RCT) of an eyecare copilot called EyeFM [1]. In this double-masked, parallel-group study, 668 participants were randomized to 16 ophthalmologists, half of whom used the EyeFM copilot and half of whom provided standard care. The primary endpoint was striking: ophthalmologists with the copilot achieved a 92.2% correct diagnostic rate compared to 75.4% in the control group — a 16.8 percentage point improvement. Referral rates also improved (92.2% vs. 80.5%), and patient compliance with self-management advice was higher in the intervention group (70.1% vs. 49.1%) at follow-up. This is the only RCT among the papers reviewed here, and it provides direct evidence that a foundation model can improve clinician performance and patient outcomes in a real-world screening setting.

How much evidence exists outside that one trial?

Outside that single RCT, the evidence base is almost entirely retrospective or laboratory-based. A systematic review of AI in perinatal medicine found that among 36 included studies, all were retrospective, single-site, or lacked external validation, and none provided prospective impact data [4]. Similarly, a systematic review of foundation models in digital pathology reported that fewer than 5% of studies were prospective, and only 14.3% conducted multicenter validation [7]. Even the largest pathology foundation model among these papers — Virchow, trained on 1.5 million whole-slide images from 100,000 patients — was evaluated on retrospective datasets, not in a prospective clinical workflow [5]. The same pattern holds for cancer imaging biomarker models: a foundation model trained on 11,467 radiographic lesions outperformed conventional methods on retrospective tasks, but no prospective trial was reported [6]. So while technical benchmarks are impressive, the gap between a model that works well on stored data and one that improves patient care in real time remains wide.

Do any models show clinical utility beyond the RCT?

A few studies provide evidence that foundation models can improve clinician efficiency or accuracy in specific tasks, even if they are not full prospective trials. For example, a literature-mining foundation model called LEADS was tested in a user study with 16 clinicians and researchers: those using LEADS achieved 0.81 recall vs. 0.78 without it, and saved 20.8% of their time in study selection [2]. For data extraction, accuracy rose from 0.80 to 0.85 with a 26.9% time saving. This is a controlled experiment, not a clinical trial, but it shows a measurable benefit in a real-world workflow. In pathology, a framework for quantifying tumor-infiltrating lymphocytes (TILs) was tested in a 'prospective deployment-style pilot' that processed slides in 2.3 minutes per slide — 87% faster than manual scoring — and showed strong concordance with expert pathologists (Pearson's r = 0.879) [3]. However, this pilot was not a randomized comparison of patient outcomes. These examples suggest that foundation models can enhance human performance in specific tasks, but the evidence is still far from proving they improve overall clinical care.

About These Sources

This answer is built on 7 peer-reviewed studies — published from 2024 to 2026, 7 from 2024 or later, 2 in Q1 journals, collectively cited 443 times — selected as the most relevant from 9 studies that passed quality screening, drawn from 74 papers retrieved from a database of over 500 million.

Sources used in this answer

1

An eyecare foundation model for clinical assistance: a randomized controlled trial.

In the only randomized controlled trial among these papers, an eyecare copilot (EyeFM) improved ophthalmologists' correct diagnostic rate from 75.4% to 92.2% and referral rate from 80.5% to 92.2% in a high-risk screening population of 668 participants.

2

A foundation model for human-AI collaboration in medical literature mining

A literature-mining foundation model (LEADS) was tested in a user study with 16 clinicians, improving recall from 0.78 to 0.81 and saving 20.8% time in study selection, and improving data extraction accuracy from 0.80 to 0.85 with 26.9% time savings.

3

Artificial intelligence based quantification of T lymphocyte infiltrate predicts prognosis in high grade breast cancer using deep learning and statistical validation

An AI framework for TIL quantification in breast cancer achieved 94.7% accuracy internally and 92.1%/91.3% on external cohorts, with a prospective pilot reducing assessment time by 87% compared to manual scoring.

4

Artificial Intelligence in Perinatal Medicine: A Systematic Review of Current Applications, Limitations, and a Translational Roadmap for the Foundation-Model Era

A systematic review of AI in perinatal medicine found that among 36 included studies, all were retrospective or single-site, with no prospective impact evaluations; the review proposed a staged roadmap for translation.

5

A foundation model for clinical-grade computational pathology and rare cancers detection

Virchow, a pathology foundation model trained on 1.5 million whole-slide images from 100,000 patients, achieved 0.95 AUC for pan-cancer detection across 16 cancer types, but was evaluated only on retrospective datasets.

6

Foundation model for cancer imaging biomarkers

A foundation model for cancer imaging biomarkers, trained on 11,467 radiographic lesions, outperformed conventional supervised methods on downstream tasks, especially with limited training data, but no prospective trial was reported.

7

Foundation Models and AI Agents in Digital Pathology Imaging: A Systematic Review of Integration into the Clinical Workflow and Implementation Challenges.

A systematic review of foundation models in digital pathology found that fewer than 5% of 42 included studies were prospective, and only 14.3% conducted multicenter validation; four implementation barriers were identified.