Researchers have developed a new token reduction schedule for Vision Transformers (ViTs) that improves performance under distribution shift. This "late-concentrated" schedule, which removes more tokens in later layers, consistently enhances out-of-distribution accuracy compared to standard "flat" schedules. The method recovers most of the original accuracy at a fraction of the compute, showing significant gains on ImageNet-C and other shift suites across various backbones and modalities without requiring per-input tuning. AI
IMPACT This research could lead to more robust and efficient Vision Transformer models, particularly in real-world applications where data distribution shifts are common.
RANK_REASON Academic paper detailing a novel method for improving model performance. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →