A new arXiv paper explores how different pretraining strategies for foundation models impact their effectiveness when transferred to ultra-widefield retinal imaging tasks. Researchers compared Vision Transformer encoders trained with supervised, Masked Autoencoder (MAE), and self-distillation objectives. Results showed that supervised and self-distillation methods outperformed MAE, with a large-scale DINOv3 model achieving the strongest performance in classifying diabetic retinopathy. The study also found that pretraining strategy influences how models aggregate evidence from different image patches, and that partial fine-tuning can improve MAE performance. AI
IMPACT Investigates how different AI model pretraining affects performance in specialized medical imaging tasks, potentially guiding future model development for healthcare applications.
RANK_REASON The cluster contains a research paper detailing experimental findings on model transferability. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- DINOv1
- DINOv3
- Mae
- Masked Autoencoder
- Mingya Alexa Gong
- vision transformer
- Vision Transformer Base
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →