How big is the gap between AI predictions and real-world materials?
AI can generate an enormous number of candidate materials, but only a tiny fraction have been experimentally verified. The largest study here, using graph neural networks trained on 48,000 known stable crystals, discovered 2.2 million new stable structures—an order-of-magnitude expansion in known materials [2]. However, of those, only 736 have been independently synthesized and confirmed in the lab [2]. That is just 0.03% of the predicted candidates. This means that for every 3,000 AI-discovered materials, only about 1 has been proven to actually exist. The rest are theoretical predictions that could be unstable, difficult to synthesize, or have different properties than expected.
A smaller but still significant study on altermagnetic materials used an AI search engine to discover 50 new candidates, covering metals, semiconductors, and insulators [1]. While these were confirmed by first-principles electronic structure calculations (a computational method), the study does not report any experimental synthesis or testing. This pattern—computational validation without lab confirmation—is common across the field and represents a key reliability risk.
How accurate are AI models, and does that guarantee real-world performance?
AI models can achieve high predictive accuracy on known data, but this does not automatically translate to reliable real-world materials. For instance, one framework for thermoelectric materials achieved over 90% accuracy in predicting properties, and identified 14 promising candidates (6 p-type, 8 n-type) that agreed with density functional theory calculations and experimental results [3]. Similarly, a machine learning model for two-dimensional topological insulators reached over 90% accuracy and discovered 56 non-trivial materials, with 17 novel insulating candidates corroborated by density functional theory [5]. These results sound impressive, but they highlight a critical nuance: the models are trained on existing data, which may not capture all possible failure modes—such as impurities, defects, or synthesis challenges—that only emerge in physical experiments.
The thermoelectric study [3] explicitly notes that obtaining effective material feature representations is still challenging and making precise predictions is tricky, despite the high accuracy. This suggests that even the best models have blind spots. Moreover, the topological insulator study [5] found that different machine learning components lead to different results, meaning the choice of algorithm and data can significantly affect which materials are flagged as promising. This variability introduces a reliability risk: two different AI systems might recommend different candidates for the same application, and neither may have been tested in a lab.
What are the practical safety and reliability risks of relying on AI-guided discovery?
The primary risk is that AI-generated materials may be unsafe or unusable in practice, leading to wasted resources or even hazardous applications. The review on machine learning in energy materials [4] points out that while AI can screen thousands of candidates quickly, the cost of first-principles calculations is still high, and the approach can get stuck in large-scale complex systems. This means that AI may flag materials that are computationally plausible but physically impossible to synthesize or that degrade under real-world conditions (e.g., temperature, humidity, stress). For example, a material predicted to be an excellent solid electrolyte for batteries might actually be toxic or react violently with other battery components—properties that AI models trained on limited data might miss.
Another risk is the 'black box' nature of many AI models. The thermoelectric study [3] and the topological insulator study [5] both rely on machine learning to learn relationships between material features and properties, often beyond human cognition [4]. While this can uncover hidden patterns, it also means that researchers may not fully understand why a material was recommended, making it harder to anticipate failure modes. The altermagnetic study [1] explicitly states that the AI search engine performed 'much better than human experts,' but this superiority is only in prediction speed and volume, not in guaranteeing real-world viability. Across all five studies, the consistent message is that AI accelerates discovery but does not eliminate the need for careful experimental validation—and the safety and reliability risks are underestimated precisely because the hype around AI's speed overshadows this critical step.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2021 to 2025, 1 from 2024 or later, 5 in Q1 journals, collectively cited 1,114 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 65 papers retrieved from a database of over 500 million.
Sources used in this answer
AI-accelerated discovery of altermagnetic materials
An AI search engine discovered 50 new altermagnetic materials (metals, semiconductors, insulators) confirmed by first-principles calculations, but no experimental synthesis was reported, highlighting a gap between prediction and real-world validation.
Scaling deep learning for materials discovery
In the largest study here, graph neural networks discovered 2.2 million stable crystal structures, but only 736 (0.03%) have been independently experimentally realized, showing that the vast majority of AI-predicted materials remain unverified.
Artificial Intelligence Guided Thermoelectric Materials Design and Discovery
An AI framework for thermoelectric materials achieved over 90% prediction accuracy and identified 14 promising candidates (6 p-type, 8 n-type) that agreed with density functional theory and experimental results, but the study notes that obtaining effective feature representations remains challenging.
Perspective on machine learning in energy material discovery
A review on machine learning in energy materials states that while AI can screen thousands of candidates quickly, first-principles calculations are still costly and can get stuck in complex systems, and AI models can learn relationships beyond human cognition, which may obscure failure modes.
Machine learning for materials discovery: Two-dimensional topological insulators
A machine learning model for two-dimensional topological insulators achieved over 90% accuracy, discovered 56 non-trivial materials (17 novel insulating candidates corroborated by density functional theory), and found that different ML components lead to different results, indicating variability in predictions.
