Vision Transformer

Vision Transformer

Description

This study developed a Vision Transformer (ViT) model with Masked Autoencoders (MAE) to classify referable diabetic retinopathy (DR) using over 100,000 large retinal images. The model, pre-trained on these retinal images, outperformed a ViT model pre-trained on ImageNet, achieving an accuracy of 93.42% and an AUC of 0.9853. The findings suggest that MAE improves classification performance and reduces the need for extensive pre-training datasets like ImageNet.

Creator

College of Science, China Jiliang University, Hangzhou, Zhejiang, China.

Information

Pediatrics or Adult

Adult

Speciality

Ophthalmology

Modality

Retinal Images

Training

Over 100,000 publicly fundus retinal images larger than 224×224

Github

Publication

FDA

Scroll to Top