Pass@1
PulseAugur coverage of Pass@1 — every cluster mentioning Pass@1 across labs, papers, and developer communities, ranked by signal.
-
New SR-PPO method improves RL for language models with single rollout
Researchers have developed a new method called Single-Rollout Proximal Policy Optimization (SR-PPO) to address the challenges of estimating token-level advantages in reinforcement learning for language models. This appr…
-
New Local Branch Routing framework enhances language model reasoning
Researchers have developed a new framework called Local Branch Routing (LBR) to improve language model reasoning during test-time scaling. LBR operates at the token level, expanding a local lookahead tree and using a li…
-
New research frames RLVR diversity collapse as overtraining
A new research paper published on arXiv explores the phenomenon of "diversity collapse" in Reinforcement Learning with Verifiable Rewards (RLVR), a technique used to enhance large language models' reasoning. The paper f…