WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

How well does AI radiology copilots handle diverse patient populations?

AI radiology copilots show promise but struggle with diverse populations due to biased training data. Performance varies by tissue, modality, and patient demographics.

Direct answer

AI radiology copilots handle diverse patient populations unevenly—they perform well on common cases but can falter on underrepresented groups. A 2024 study found that a top AI pathology copilot (PathChat) achieved state-of-the-art accuracy on diagnostic questions across diverse tissue types and diseases [1], but a separate 2021 analysis warns that most biomedical AI is trained on non-representative samples, risking biased outcomes for minority populations [3]. The bottom line: these tools work best when trained on broad, inclusive data, and they still need human oversight, especially for patients outside the training set.

4sources cited

This article was generated with WisPaper-powered search and paper analysis.

How well do AI radiology copilots perform across different patient groups?

The strongest evidence here shows that AI copilots can perform well across a range of tissue types and diseases, but their success depends heavily on the diversity of the data they were trained on. In a 2024 study, the PathChat AI copilot was trained on over 456,000 visual-language instructions covering diverse tissue origins and disease models, and it achieved state-of-the-art performance on multiple-choice diagnostic questions [1]. That means it handled questions about many different kinds of tissue and disease better than other AI systems. However, the same study notes that the system was fine-tuned specifically for pathology, so its performance on radiology images (like CT or MRI) wasn't directly tested [1].

A 2021 analysis of biomedical AI broadly warns that algorithms are often developed on non-representative samples, which can lead to biased results for groups not well represented in the training data [3]. This means that if an AI radiology copilot is trained mostly on images from one ethnic group or one type of hospital, it may be less accurate for patients from other backgrounds. The authors recommend collecting more diverse data and monitoring AI performance across subgroups to catch disparities [3].

What about different imaging modalities and body parts?

AI copilots are being developed to handle multiple imaging types, but the evidence shows that performance varies by modality and body region. A 2024 project produced AI-generated annotations for 11 different cancer image collections, including CT, MRI, and PET scans, covering body parts like the chest, breast, kidneys, prostate, and liver [2]. The fact that the AI could generate segmentations across these varied modalities and body parts suggests it can handle diverse inputs. However, the same study notes that only 4% of the original DICOM studies had existing segmentation annotations, meaning the AI was working from a very sparse starting point [2]. A radiologist reviewed and corrected a portion of the AI's annotations to assess performance, implying that the AI's raw output wasn't always accurate enough to use without human checking [2].

This aligns with the broader concern from the 2021 analysis: even when AI works across multiple modalities, its accuracy can be uneven if the training data doesn't adequately represent all the variations it will encounter in practice [3].

Can patients trust AI diagnostic reports equally, regardless of background?

Patient trust in AI diagnostic reports is not uniform—it depends on the patient's health literacy and how the AI explains its findings. A 2023 study found that patients with different levels of health literacy had different levels of trust in AI diagnostic reports, and they also looked at the reports differently (as measured by eye-tracking) [4]. Specifically, the type of explanation the AI gave (global vs. partial) affected trust, and the effect varied by the patient's health literacy [4]. This means that even if the AI copilot is technically accurate, patients from different backgrounds may not trust it equally, which could affect how they act on the results. The study suggests that AI diagnostic tools need to tailor their explanations to the user to build appropriate trust [4].

This finding adds a human dimension to the technical challenges: diverse patient populations bring diverse expectations and understanding, which AI copilots must account for to be truly useful.

About These Sources

This answer is built on 4 peer-reviewed studies — published from 2021 to 2024, 2 from 2024 or later, 2 in Q1 journals, collectively cited 411 times — selected as the most relevant from 4 studies that passed quality screening, drawn from 48 papers retrieved from a database of over 500 million.

Sources used in this answer

1

A multimodal generative AI copilot for human pathology

PathChat, a vision-language AI copilot for pathology, was trained on over 456,000 diverse visual-language instructions and achieved state-of-the-art performance on multiple-choice diagnostic questions across diverse tissue origins and disease models, though it was not tested on radiology images specifically.

2

AI-Generated Annotations Dataset for Diverse Cancer Radiology Collections in NCI Image Data Commons.

AI-generated annotations for 11 cancer image collections (CT, MRI, PET) covering chest, breast, kidneys, prostate, and liver were produced, but only 4% of original studies had existing segmentations, and radiologist review was needed to correct some AI outputs.

3

Ensuring that biomedical AI benefits diverse populations

Biomedical AI algorithms are often developed on non-representative samples, which can introduce bias and health disparities; the authors recommend more diverse data collection and monitoring across subgroups.

4

Impact and Prediction of AI Diagnostic Report Interpretation Type on Patient Trust

Patient trust in AI diagnostic reports varies by health literacy and the type of explanation provided (global vs. partial), as measured by eye-tracking and surveys, indicating that AI copilots must tailor explanations to diverse users.