Why does poor data quality cripple AI drug repurposing?
AI models are only as good as the data they are trained on, and current drug repurposing systems suffer from data that is incomplete, biased, or poorly standardized. A 2024 review notes that data quality issues—such as missing clinical records, inconsistent coding in electronic health records (EHRs), and sparse genomic information—directly limit the reliability of AI predictions [3]. For instance, a study using EHR data from 12,253 MGUS patients had to exclude nearly 4,500 patients because medication data was unavailable, and even then, the model's hazard ratios for drug effects were derived from bootstrapping rather than direct clinical trials [4]. This means many AI-identified candidates may be artifacts of data gaps rather than true therapeutic signals.
The problem is compounded by the narrow focus of existing datasets. A 2021 expert review points out that most anticancer drug repurposing approaches rely heavily on genomics-based information, which misses functional drug responses that can only be captured by testing drugs on actual patient cells [5]. This scarcity of clinical patient data means AI models often learn from indirect associations rather than direct evidence of efficacy, leading to high false-positive rates. Across the papers here, the consensus is clear: without higher-quality, more comprehensive training data—including functional testing results and complete EHRs—AI systems will continue to produce predictions that are statistically promising but clinically unreliable [3][5][6].
How does ignoring sex differences undermine AI predictions?
A major but often overlooked evidence gap is the failure to account for sex differences in drug response, which can render AI predictions inaccurate for half the population. A 2022 review highlights that sex differences—including hormonal, immunological, and metabolic variations—cause many drugs to work differently in males and females, yet most drug repurposing methods are 'sex-blind' [2]. For example, the same drug might be more efficacious in one sex or cause sex-biased adverse events (SBAEs) in the other, but current AI models rarely incorporate this information because training datasets are often imbalanced, with disproportionate numbers of male or female samples [2]. This imbalance leads to poorer model performance for the underrepresented sex.
The review further notes that while sex-aware methods exist for clinical, genomic, and transcriptomic data, they have not expanded to other data types like DNA variation, which has proven useful in non-sex-aware repurposing [2]. This gap means AI systems may miss repurposing opportunities that are sex-specific or, worse, recommend drugs that cause harm in one sex. The authors argue that low-dimensional representations of molecular association and network approaches are promising for future sex-aware methods, but the lack of large, balanced datasets remains a critical barrier [2]. No other paper in this set addresses sex differences, making this a unique and urgent gap.
Why do AI predictions still need real-world testing?
Even the most advanced AI models for drug repurposing lack robust clinical validation, meaning their predictions are hypotheses, not proven therapies. A 2024 study introduced TxGNN, a graph foundation model that achieved a 49.2% improvement in prediction accuracy for drug indications and 35.1% for contraindications under zero-shot conditions—meaning it could predict uses for diseases with no existing treatments [1]. While impressive, the authors note that many of these predictions align with off-label prescriptions already made by clinicians, suggesting the model is rediscovering known uses rather than reliably identifying novel ones [1]. The model's 'Explainer' module performed well in human evaluations, but these were assessments of interpretability, not clinical outcomes.
Across the literature, the call for clinical validation is consistent. A 2024 review emphasizes that AI-driven repurposing faces regulatory hurdles and requires interdisciplinary collaboration to translate predictions into clinical trials [6]. Similarly, a 2023 study using machine learning on EHR data for MGUS progression explicitly states that its findings 'could inform subsequent prospective studies'—meaning the results are preliminary and need confirmation in controlled trials [4]. The 2021 review on cancer drug repurposing adds that functional testing of patient cells exposed to drugs is an essential validation step that AI models currently bypass [5]. In short, without prospective clinical trials or at least rigorous real-world evidence, AI predictions remain educated guesses, and the field's most advanced models still fall short of the evidence needed for regulatory approval or clinical adoption [1][4][6].
About These Sources
This answer is built on 6 peer-reviewed studies — published from 2021 to 2025, 3 from 2024 or later, 4 in Q1 journals, collectively cited 408 times — selected as the most relevant from 9 studies that passed quality screening, drawn from 48 papers retrieved from a database of over 500 million.
Sources used in this answer
A foundation model for clinician-centered drug repurposing
Introduces TxGNN, a graph foundation model that improves zero-shot prediction accuracy for drug indications by 49.2% and contraindications by 35.1% across 17,080 diseases, but its predictions align with off-label use rather than novel discoveries.
Considerations and challenges for sex-aware drug repurposing
Highlights that sex-aware drug repurposing methods are underdeveloped due to imbalanced training data and lack of expansion to data types like DNA variation, limiting their ability to predict sex-specific drug responses.
Artificial Intelligence-Based Methods for Drug Repurposing and Development in Cancer
Reviews AI methods for cancer drug repurposing, noting that data quality, interpretability, and ethical considerations are major challenges limiting clinical translation.
Artificial intelligence-enabled screening strategy for drug repurposing in monoclonal gammopathy of undetermined significance
Uses machine learning on EHR data from 12,253 MGUS patients to identify drug classes associated with lower progression risk, but results are preliminary and require prospective validation.
Artificial intelligence, machine learning, and drug repurposing in cancer
Argues that scarcity of clinical patient data and over-reliance on genomics limit anticancer drug repurposing, and recommends functional testing of patient cells as an additional validation source.
Artificial Intelligence in Clinical Practice: Unlocking New Horizons in Drug Repurposing for Disease Treatment
Reviews AI applications in drug repurposing for rare diseases, noting that data quality, regulatory hurdles, and need for interdisciplinary collaboration are key barriers.
