This study developed a Vision Transformer (ViT) model with Masked Autoencoders (MAE) to classify referable diabetic retinopathy (DR) using over 100,000 large retinal images. The model, pre-trained on these retinal images, outperformed a ViT model pre-trained on ImageNet, achieving an accuracy of 93.42% and an AUC of 0.9853. The findings suggest that MAE improves classification performance and reduces the need for extensive pre-training datasets like ImageNet.
Creator
College of Science, China Jiliang University, Hangzhou, Zhejiang, China.
Information
Pediatrics or Adult
Adult
Speciality
Ophthalmology
Modality
Retinal Images
Training
Over 100,000 publicly fundus retinal images larger than 224×224