Researchers have introduced OptiSelect, a new framework for optimizing data selection during large language model pretraining. The framework formalizes the concept of optimizer-aware selection, which considers how the optimizer step influences the value of data candidates. Theoretical analysis suggests that diagonal-adaptive optimizers like AdamW and Sophia are more effective for this selection process than sign-based or polar-tangential preconditioners such as Lion and Muon. Experimental results with 124M and 720M models support these findings, indicating that AdamW's scoring geometry performs best even when Muon is used as the optimizer. AI
IMPACT Provides theoretical foundations and practical guidance for co-designing optimizers and data selection in LLM pretraining.
RANK_REASON The cluster contains a research paper detailing a new framework and theoretical analysis for optimizing LLM pretraining. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →