Researchers have identified a significant geometric bottleneck in knowledge distillation between Vision Transformers and smaller CNNs. Standard cosine distillation causes the learned representations to collapse to a low dimensional space, regardless of the CNN's parameter count. While an auxiliary InfoNCE objective can expand this dimensionality, it paradoxically degrades downstream accuracy by 15-18 points. A label-aware Supervised Contrastive distillation method, however, shows promise by increasing dimensionality while maintaining or improving accuracy, suggesting that the utility of dimensionality expansion depends on the objective being label-aware. AI
IMPACT Highlights the importance of label-aware objectives in cross-modal distillation for effective representation learning.
RANK_REASON Academic paper detailing novel findings in AI model distillation techniques. [lever_c_demoted from research: ic=1 ai=1.0]
- CIFAR-10
- CIFAR-100
- CLIP ViT-B/32
- CNNs
- InfoNCE
- Kabir Thayani
- Supervised Contrastive distillation
- Vision Transformers
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →