A Comparison of Machine Learning and Human Graders for Glaucoma Diagnosis from Fundus Images for Population Screening
- PMID: 42575300
- DOI: 10.1016/j.ophtha.2026.07.043
A Comparison of Machine Learning and Human Graders for Glaucoma Diagnosis from Fundus Images for Population Screening
Abstract
Purpose: To compare the accuracy of vertical cup-disc ratios (VCDRs), ascertained by machine learning (ML) versus human graders, from fundus images for glaucoma detection. This study uses population-based data, with a disease prevalence and case mix that is more reflective of an unselected patient population than conventional case-control studies, with the aim of developing improved glaucoma screening tests.
Design: Cross-sectional analysis of a population-based study.
Participants: A total of 6304 participants of the European Prospective Investigation of Cancer (EPIC)-Norfolk Eye Study with color fundus images gradable by humans and ML in both eyes.
Methods: Vertical cup-disc ratio was independently estimated from 2-dimensional fundus images of EPIC-Norfolk Eye Study participants by trained human graders (human-derived VCDR [H-VCDR]) and an externally trained, open-access ML model (ML-derived VCDR [ML-VCDR]). A neural network trained on 81 830 ophthalmologist-labeled images was used to generate pseudo-labels for over 100 000 UK Biobank images, on which ML-VCDR was subsequently trained. Glaucoma status was ascertained by tertiary-center specialist examination. Predictive performance of VCDR for glaucoma status was examined using logistic regression. Machine learning-derived VCDR estimates were additionally compared with a popular open-source ML model (AutoMorph) and scanning laser ophthalmoscopy (Heidelberg Retinal Tomography [HRT]).
Main outcome measures: Area under the receiver operating characteristic curve (AUROC) explained variance (McFadden's pseudo-R2).
Results: Of 6304 participants (mean age, 68 years; 57% women), 696 had glaucoma or suspect status in at least 1 eye. For left eyes, H-VCDR and ML-VCDR explained 17% (95% confidence interval [CI], 14.7-20.4) and 31% (95% CI, 27.9-33.9) of glaucoma status variance and had an AUROC of 79% (95% CI, 76.8-81.2) and 88% (95% CI, 86.3-89.1), respectively. Right eye H-VCDR and ML-VCDR explained 20% (95% CI, 16.9-22.6) and 35% (95% CI, 32.4-37.9) of the variance and had an AUROC of 81% (95% CI, 78.6-82.5) and 90% (95% CI, 88.7-91.0), respectively. Machine learning-derived VCDR also performed better than AutoMorph and HRT at predicting glaucoma status from VCDR estimations.
Conclusions: In this population-based setting, ML far outperformed trained human graders at predicting specialist-ascertained glaucoma status from fundus images. This provides promise for ML-supported strategies for glaucoma population screening.
Financial disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Keywords: Glaucoma; Machine learning; Screening; VCDR.
Copyright © 2026 American Academy of Ophthalmology, Inc. Published by Elsevier Inc. All rights reserved.
LinkOut - more resources
Full Text Sources
