ENTITY
language model pre-training
language model pre-training
PulseAugur coverage of language model pre-training — every cluster mentioning language model pre-training across labs, papers, and developer communities, ranked by signal.
Total · 30d
2
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D
1 day(s) with sentiment data
RECENT · PAGE 1/1 · 2 TOTAL
-
New research explores language model pre-training dynamics and proposes tuning strategies
A new research paper analyzes the pre-training dynamics of language models from the perspective of local landscape geometry. The study identifies two distinct phases: Phase I, where sharpness leads to instability with l…
-
New TBP parameterizations enhance hyper-connection expressivity and stability
Researchers have introduced Transportation Birkhoff Polytope (TBP) parameterizations as a novel method for constructing exactly doubly stochastic mixing matrices in hyper-connections. This approach offers full expressiv…