A new research paper explores the stability dynamics of Zeroth-Order (ZO) optimization methods, particularly in the context of deep learning. The study identifies a specific step size condition that governs the linear stability of these methods, contrasting them with First-Order (FO) methods. Unlike FO methods, which are influenced by the largest Hessian eigenvalue, ZO methods' stability depends on the entire Hessian spectrum. The research also proposes tractable stability bounds based on the largest eigenvalue and Hessian trace, finding that full-batch ZO methods operate at the edge of stability in deep learning tasks. AI
IMPACT Provides theoretical insights into optimization techniques relevant for memory-efficient fine-tuning of large models.
RANK_REASON Academic paper on optimization methods in machine learning. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →