Alexandru Meterez
PulseAugur coverage of Alexandru Meterez — every cluster mentioning Alexandru Meterez across labs, papers, and developer communities, ranked by signal.
-
Anytime Pretraining offers horizon-free LLM training with weight averaging
Researchers have introduced "Anytime Pretraining," a novel approach to training large language models that eliminates the need for pre-defined training horizons. This method utilizes horizon-free learning-rate schedules…
-
Seesaw method accelerates LLM training by optimizing batch size and learning rate
Researchers have developed a new method called Seesaw to accelerate the training of large language models by optimizing the scheduling of batch sizes and learning rates. This approach theoretically demonstrates an equiv…
-
Quadratic Model Shows Predictive Power for LLM Optimization Dynamics
A new paper published on arXiv proposes that the simple quadratic model can be a surprisingly accurate predictor of optimization dynamics in large language models (LLMs). Researchers demonstrated that by analyzing the H…