Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2026 Aug 10:S0161-6420(26)00565-8.
doi: 10.1016/j.ophtha.2026.07.043. Online ahead of print.

A Comparison of Machine Learning and Human Graders for Glaucoma Diagnosis from Fundus Images for Population Screening

Affiliations
Free article

A Comparison of Machine Learning and Human Graders for Glaucoma Diagnosis from Fundus Images for Population Screening

Thomas R P Taylor et al. Ophthalmology. .
Free article

Abstract

Purpose: To compare the accuracy of vertical cup-disc ratios (VCDRs), ascertained by machine learning (ML) versus human graders, from fundus images for glaucoma detection. This study uses population-based data, with a disease prevalence and case mix that is more reflective of an unselected patient population than conventional case-control studies, with the aim of developing improved glaucoma screening tests.

Design: Cross-sectional analysis of a population-based study.

Participants: A total of 6304 participants of the European Prospective Investigation of Cancer (EPIC)-Norfolk Eye Study with color fundus images gradable by humans and ML in both eyes.

Methods: Vertical cup-disc ratio was independently estimated from 2-dimensional fundus images of EPIC-Norfolk Eye Study participants by trained human graders (human-derived VCDR [H-VCDR]) and an externally trained, open-access ML model (ML-derived VCDR [ML-VCDR]). A neural network trained on 81 830 ophthalmologist-labeled images was used to generate pseudo-labels for over 100 000 UK Biobank images, on which ML-VCDR was subsequently trained. Glaucoma status was ascertained by tertiary-center specialist examination. Predictive performance of VCDR for glaucoma status was examined using logistic regression. Machine learning-derived VCDR estimates were additionally compared with a popular open-source ML model (AutoMorph) and scanning laser ophthalmoscopy (Heidelberg Retinal Tomography [HRT]).

Main outcome measures: Area under the receiver operating characteristic curve (AUROC) explained variance (McFadden's pseudo-R2).

Results: Of 6304 participants (mean age, 68 years; 57% women), 696 had glaucoma or suspect status in at least 1 eye. For left eyes, H-VCDR and ML-VCDR explained 17% (95% confidence interval [CI], 14.7-20.4) and 31% (95% CI, 27.9-33.9) of glaucoma status variance and had an AUROC of 79% (95% CI, 76.8-81.2) and 88% (95% CI, 86.3-89.1), respectively. Right eye H-VCDR and ML-VCDR explained 20% (95% CI, 16.9-22.6) and 35% (95% CI, 32.4-37.9) of the variance and had an AUROC of 81% (95% CI, 78.6-82.5) and 90% (95% CI, 88.7-91.0), respectively. Machine learning-derived VCDR also performed better than AutoMorph and HRT at predicting glaucoma status from VCDR estimations.

Conclusions: In this population-based setting, ML far outperformed trained human graders at predicting specialist-ascertained glaucoma status from fundus images. This provides promise for ML-supported strategies for glaucoma population screening.

Financial disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Keywords: Glaucoma; Machine learning; Screening; VCDR.

PubMed Disclaimer

LinkOut - more resources