Researchers have developed a new knowledge distillation technique called Anti-Shortcut Distillation (ASD). This method uses an early-stage teacher model as a negative reference to guide a student model away from learning shortcuts. ASD incorporates two losses: a temporal contrastive loss and a shortcut suppression loss, which penalizes the student's projection onto identified shortcut directions. Experiments on CIFAR-100, ImageNet-100, and TinyImageNet show ASD outperforms standard knowledge distillation in clean accuracy and corruption robustness, particularly in cross-architecture scenarios. AI
IMPACT Introduces a novel technique to improve model robustness by explicitly teaching models to avoid shortcut learning.
RANK_REASON This is a research paper detailing a novel method for knowledge distillation. [lever_c_demoted from research: ic=1 ai=1.0]
- Anti-Shortcut Distillation
- CIFAR-100
- ImageNet-100
- knowledge distillation
- ShuffleNet-V2
- Syed Muhammad Raza
- TinyImageNet
- WRN-40-2
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →