A new arXiv paper reveals that increasing the size of text encoders in CLIP models can negatively impact zero-shot performance, even when the total parameter count increases. Researchers found that for most vision encoders, there's an optimal text encoder size beyond which performance degrades due to overfitting. The study suggests that modality-specific weight decay coefficients can recover and improve performance in these degraded configurations. The findings highlight a trade-off between embedding uniformity and cross-modal alignment, which are predictive of zero-shot performance, and aim to guide more efficient and reliable scaling of CLIP architectures. AI
IMPACT Findings may lead to more efficient CLIP model training and improved zero-shot capabilities.
RANK_REASON Academic paper detailing research findings on model architecture. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →