A new benchmark study revisits the challenge of predicting deep learning model generalization before target-test evaluation, focusing on settings with controlled covariate shift rather than independent and identically distributed (IID) data. Researchers found that the effectiveness of generalization measures is highly dependent on the specific regime, such as image classifiers evaluated under various corruptions and perturbations like those in CIFAR-10-C/P. The study suggests that model selection should not rely solely on measures favored by IID evaluations, but rather treat generalization measures as regime-dependent ranking signals whose utility must be assessed for the intended corruption or perturbation setting. AI
IMPACT Findings suggest that model selection strategies need to adapt to specific data corruption or perturbation settings, rather than relying on IID-based measures.
RANK_REASON Academic paper introducing a new benchmark and findings on generalization measures in deep learning. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX Code Finder for Papers
- CIFAR-10-C/P
- CORE Recommender
- DagsHub
- Dziugaite et al.
- Gotit.pub
- Hugging Face
- Influence Flower
- Jiang et al.
- ScienceCast
- Sora Nakai
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →