A new paper published on arXiv details the training dynamics of multiclass logistic regression, establishing precise scaling laws for cross-entropy risk under gradient-based optimization. The research indicates that learning progresses sequentially across classes, from most to least frequent, and identifies three distinct phases in this process: an initial plateau, a power-law decay regime, and a final convergence regime. The study also analyzes the interplay between model capacity and optimization under a fixed compute budget, deriving a compute-optimal scaling law for logistic regression that prescribes model size and training time based on available compute. AI
RANK_REASON Academic paper detailing theoretical scaling laws for a machine learning model. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →