The vertical cup-to-disc ratio, measured from the optic nerve head, is a key structural marker in glaucoma screening. This capstone, sponsored by UVA Ophthalmology, built a reproducible pipeline that segments the optic disc and cup from retinal fundus images and estimates that ratio, then measured how well a model trained on public data works on real clinic images. It received the Most Innovative Analytical Solution award.
Team project with Robert Judson Ashby, Emmanuel Gyamfi and Michael Ieraci. Sponsor: Dr. Arjun Dirghangi. Faculty mentor: Dr. Aiying Zhang.
The public data is 3,358 fundus images with disc and cup annotations from four datasets: ORIGA (650), G1020 (1,020), REFUGE (1,200) and PAPILA (488). ORIGA, G1020 and PAPILA use leakage-aware, group-wise 70/15/15 splits, and REFUGE keeps its official partitions.
The clinical data is de-identified fundus images from the UVA Department of Ophthalmology, accessed under a data-use agreement with required human-subjects training. Ground truth came from clinically provided annotations, giving 59 mask-ready samples from 20 patient or encounter groups. That data is private and isn't included in the repository.
We compared U-Net, U-Net++ and DeepLabV3+ under a shared training budget, with images resized to 256 by 256, encoders trained from scratch and a combined Dice and cross-entropy loss. U-Net++ with a ResNet-18 encoder was selected.
With the architecture fixed, we varied the data pipeline instead of the model: online augmentation, a synthetic expansion strategy for probing data scaling, and a longer 25-epoch schedule. Extended training gave the clearest public gain, raising held-out mean foreground Dice from 0.818 to 0.842.
| Evaluation | Metric | Value |
|---|---|---|
| Public-only model, long training | Public test mean foreground Dice | 0.842 |
| Zero-shot on clinical images | Patient-weighted Dice | 0.251 |
| Hybrid model | Public test mean foreground Dice | 0.844 |
| Hybrid clinical adaptation | Patient-weighted Dice | 0.265 → 0.330 |
| Hybrid clinical adaptation | Cup-to-disc ratio error reduction | 0.122 |
The system is exploratory and not clinically deployable. It reports structural measurements — segmentation and cup-to-disc ratio — not a diagnosis. The clinical set is small, and the disc Dice definition isn't yet consistent between public and clinical evaluation.
The highest-value next step is a larger, segmentation-ready clinical dataset. Others include ImageNet-pretrained encoders with matched normalisation, and moving from cup-to-disc ratio toward automated DDLS scoring and from still frames toward video ophthalmoscopy.