WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

Can artificial intelligence detect diabetic retinopathy accurately?

AI detects diabetic retinopathy with high accuracy, often matching or exceeding specialists, but performance varies by setting and image quality.

Direct answer

Yes, artificial intelligence can detect diabetic retinopathy accurately, often matching or exceeding the performance of human specialists. In large clinical trials, AI systems have shown sensitivity (ability to correctly identify disease) of 95-100% and specificity (ability to correctly rule out disease) of 85-98% [2][3][5]. However, accuracy drops in real-world settings with poor image quality, cataracts, or when using portable cameras [9][11]. Across the studies here, the largest and most rigorous trials consistently show AI is a reliable screening tool, but it is not perfect and works best as a complement to, not a replacement for, specialist review.

11sources cited

This article was generated with WisPaper-powered search and paper analysis.

How accurate is AI compared to human experts?

In head-to-head comparisons, AI systems often detect diabetic retinopathy (DR) more sensitively than human graders, meaning they catch more cases. A pivotal multicenter trial of the EyeArt system found it detected more-than-mild DR with 95.5% sensitivity and 85.0% specificity, while general ophthalmologists detected only 20.6% of cases and retina specialists 59.5% [3]. Another study showed AI outperformed a retina specialist in detecting vision-threatening DR (96.2% vs 84.4% sensitivity) [7]. This means AI is less likely to miss disease, though it may flag more false positives.

However, specificity (correctly ruling out disease) is often slightly lower for AI than for specialists. In the same EyeArt trial, AI had 86-88% specificity compared to 98.9-99.8% for human graders [3]. A study in rural China found AI had 81.2% sensitivity and 94.3% specificity relative to ophthalmologists [6]. So while AI catches more true cases, it also generates more false alarms that require specialist follow-up.

What affects AI accuracy in real-world use?

Image quality is the biggest factor. A study simulating cataracts found AI accuracy dropped from 97.0% to 53.5% as cataract severity increased [9]. Similarly, using a handheld portable fundus camera instead of a standard desktop camera reduced AI's ability to detect referable DR (area under the curve fell from 98.5% to 89.4%) [11]. In community screenings, cataracts and other media opacities reduced accuracy [8].

The type of AI system and the population also matter. In a direct comparison of two FDA-cleared systems, IDx-DR had higher sensitivity (99.1%) but lower specificity (71.5%) for referable DR, while MONA DR had lower sensitivity (93.4%) but higher specificity (89.3%) [4]. A study in Indigenous Australians found AI maintained high sensitivity (98%) and specificity (95%) for more-than-mild DR, showing it can work well across ethnic groups [7]. However, a Chilean study warned that reported accuracy can be inflated if the wrong reference standard is used [10].

Can AI replace human screeners?

AI is best seen as a powerful screening tool, not a replacement for specialists. Multiple studies emphasize that AI's strength is in ruling out disease—its negative predictive value (NPV) is often 100%, meaning a negative result is highly reliable [2][8]. This allows AI to safely reduce the number of patients needing specialist review. For example, one study found AI had 100% NPV for all DR stages, so clinicians can trust a 'no DR' result [2].

However, AI tends to overestimate disease severity, so all positive results need specialist confirmation [2]. In a community screening, EyeArt had 81% accuracy for diagnosis and 83% accuracy for referrals compared to a retina specialist [8]. Successful implementation also requires proper clinic workflow, staff training, and integration with referral systems [1]. AI can improve access and equity, especially in underserved areas [1][6][7], but it is not yet a standalone solution.

About These Sources

This answer is built on 11 peer-reviewed studies — published from 2021 to 2025, 5 from 2024 or later, 4 in Q1 journals, collectively cited 386 times — selected as the most relevant from 15 studies that passed quality screening, drawn from 87 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Autonomous Artificial Intelligence in Diabetic Retinopathy Testing—Lessons Learned on Successful Health System Adoption

Reviews 5 years of real-world adoption of FDA-cleared AI systems (LumineticsCore, EyeArt, AEYE-DS), reporting average sensitivity 87-100%, specificity 60-91%, and highlighting successful implementation requires workflow integration and staff training.

2

Accuracy of Autonomous Artificial Intelligence-Based Diabetic Retinopathy Screening in Real-Life Clinical Practice

In a study of 1141 patients, the IDx-DR AI system had 100% sensitivity for all DR stages and 100% negative predictive value, but overestimated severity, so positive results need specialist review.

3

Artificial Intelligence Detection of Diabetic Retinopathy

In a prospective multicenter trial of 893 participants, EyeArt AI detected more-than-mild DR with 95.5% sensitivity and 85% specificity, outperforming general ophthalmologists (20.6% sensitivity) and retina specialists (59.5% sensitivity).

4

Evaluating the efficacy of AI systems in diabetic retinopathy detection: A comparative analysis of Mona DR and IDx-DR.

Comparing two AI systems in 695 patients, IDx-DR had 99.1% sensitivity and 71.5% specificity for referable DR, while MONA DR had 93.4% sensitivity and 89.3% specificity; both had 100% sensitivity for sight-threatening DR.

5

Pivotal Evaluation of an Artificial Intelligence System for Autonomous Detection of Referrable and Vision-Threatening Diabetic Retinopathy

In a pivotal multicenter study of 893 patients, EyeArt AI detected more-than-mild DR with 95.5% sensitivity and 85% specificity, and vision-threatening DR with 95.1% sensitivity and 89% specificity without dilation.

6

Clinical evaluation of AI-assisted screening for diabetic retinopathy in rural areas of midwest China

In a rural Chinese screening of 3933 patients, AI had 81.2% sensitivity and 94.3% specificity compared to ophthalmologists, with high consistency (kappa=0.752).

7

Validation of a deep learning system for the detection of diabetic retinopathy in Indigenous Australians.

In 864 Indigenous Australian patients, a deep learning system had superior sensitivity (98%) to a retina specialist (87.1%) for more-than-mild DR, with slightly lower specificity (95.1% vs 97%).

8

EyeArt artificial intelligence analysis of diabetic retinopathy in retinal screening events

In community screenings (124 eyes), EyeArt had 81% accuracy for DR diagnosis and 83% accuracy for referrals compared to a retina specialist, with substantial agreement (kappa=0.69-0.70).

9

Effect of simulated cataract on the accuracy of artificial intelligence in detecting diabetic retinopathy in color fundus photos.

Simulated cataracts reduced AI accuracy from 97% to 53.5%, with sensitivity dropping from 95.7% to 11.8%, showing image quality is critical for AI performance.

10

Amending AI Software Accuracy for Diabetic Retinopathy Detection Using Conditional Probability and the Appropriate Reference Standard.

Using conditional probability to correct for an inappropriate reference standard, the reported sensitivity and specificity of the DART AI system were significantly lower than originally claimed.

11

Evaluation of an AI system for the detection of diabetic retinopathy from images captured with a handheld portable fundus camera: the MAILOR AI study.

An AI system (Pegasus) had good performance for proliferative DR (AUROC 94.3%) but significantly lower performance for referable DR (AUROC 89.4%) when using a handheld portable fundus camera compared to a desktop camera.