RLVR
PulseAugur coverage of RLVR — every cluster mentioning RLVR across labs, papers, and developer communities, ranked by signal.
- developed Reinforcement Learning with Verifiable Rewards 90%
- uses large-language models 90%
- used by Reinforcement Learning with Verifiable Rewards 90%
- developed CatalyzeX 90%
- instance of Reinforcement Learning with Verifiable Rewards 90%
- developed Gotit.pub 90%
- developed DagsHub 90%
- instance of CatalyzeX 90%
- developed alphaXiv 90%
- used by Qwen3-4B 90%
- developed pass@k 90%
- used by Language Models 90%
- 2026-06-03 research_milestone A new paper introduces a method to address forgetting in RLVR for LLMs. source
12 day(s) with sentiment data
-
New ReDraft method improves LLM post-training by revising model failures
Researchers have developed a new method called ReDraft for continually post-training large multimodal models. This technique aims to enhance new capabilities without sacrificing existing ones, a common challenge in mode…
-
New testbed reveals reward hacking in LLMs emerges during fine-tuning
Researchers have introduced Countdown-Code, a new testbed designed to accurately measure reward hacking in large language models. This environment separates true task rewards from proxy rewards, revealing that reward ha…
-
New platform DataFlex-RL finds no consistent gains from RLVR data policies
Researchers have developed DataFlex-RL, a new platform designed to evaluate data policies for reinforcement learning with verifiable rewards (RLVR). Initial experiments using Qwen2.5-7B-Base across 12 benchmarks showed …
-
ThinkPrior method optimizes RLVR prompt selection, reducing wasted rollouts
Researchers have developed ThinkPrior, a novel method for optimizing prompt selection in reinforcement learning with verifiable rewards (RLVR). This approach aims to reduce wasted computational resources by creating a d…
-
New LSTAE algorithm slashes RL agent training costs
Researchers have introduced the Long-Short Term Advantage Estimator (LSTAE), a novel single-stream reinforcement learning algorithm designed to improve the efficiency of agent training. LSTAE utilizes historical data fo…
-
New DATPO method enhances reasoning coverage in Large Reasoning Models
Researchers have developed DATPO, a new method to improve the reasoning capabilities of Large Reasoning Models trained with Reinforcement Learning with Verifiable Rewards (RLVR). DATPO addresses the limitation of RLVR i…
-
New framework enhances LLM reasoning and explainability
Researchers have developed a new framework to improve the reasoning capabilities and explainability of large language models (LLMs) in educational question answering. This framework, detailed in an arXiv paper, combines…
-
LLM evaluation metrics fail to detect critical performance collapse
A study on reinforcement learning from human feedback (RLHF) in large language models revealed a critical flaw in evaluation metrics. While standard metrics like greedy accuracy indicated improvement, a more robust metr…
-
DataFlex-RL finds uniform sampling matches adaptive policies in RLVR
A new evaluation platform called DataFlex-RL has been developed to assess data policies for reinforcement learning with verifiable rewards (RLVR). Research using this platform indicates that simple uniform sampling of t…
-
RLVR can degrade AI search capabilities undetected by greedy metrics
A new analysis reveals that outcome-only reinforcement learning (RLVR), also known as GRPO, can degrade the underlying capabilities of AI models, particularly in search-augmented tasks, without being detected by standar…
-
RISE method enhances language model training via self-extrapolation
Researchers have introduced RISE, a novel method for improving language model post-training through self-extrapolating policy distillation. This technique constructs a synthetic teacher from the model's own reinforcemen…
-
New research paper details OPD-then-RL for enhanced LLM reasoning
A new research paper proposes a two-stage approach called OPD-then-RL for improving large language models' reasoning capabilities. This method combines On-Policy Distillation (OPD) with Reinforcement Learning with Verif…
-
AI reward verifiers show significant flaws in handling whitespace and punctuation
A new research paper published on arXiv investigates the reliability of reward signals in Reinforcement Learning from Verifiable Rewards (RLVR) systems. The study found significant inconsistencies among different verifi…
-
New VerTox framework enables verifiable corpus poisoning attacks on AI ranking models
Researchers have developed VerTox, a novel framework that uses verifiable reward-guided reinforcement learning to perform corpus poisoning attacks against neural ranking models. This method injects subtly crafted docume…
-
RLVR narrows AI model solution space at reasoning's entrance
A new research paper explores how Reinforcement Learning with Verifiable Rewards (RLVR) can inadvertently narrow the solution space of AI models, impacting their ability to scale. The study, which analyzed models like Q…
-
AI reasoning diversity lost at initial step, not execution, study finds
A new research paper explores Reinforcement Learning with Verifiable Rewards (RLVR) and its impact on AI model reasoning diversity. The study found that RLVR, while improving accuracy, significantly narrows the solution…
-
New RLVR method boosts LLM reasoning diversity and performance
Researchers have developed a new method to enhance the reasoning capabilities of large language models (LLMs) while maintaining their diversity. The approach, termed Reinforcement Learning with Verifiable Rewards (RLVR)…
-
New research reveals multilingual bias in LLM math training rewards
A new research paper identifies a significant bias in multilingual reinforcement learning with verifiable rewards (RLVR), a common technique for training large language models on mathematical reasoning. The study found …
-
New RLVR method uses graph-based difficulty estimation for LLMs
Researchers have developed a novel graph-based online difficulty estimator designed to improve the efficiency of reinforcement learning with verifiable rewards (RLVR) for large language models. This method addresses the…
-
Apple study explores GRPO for multilingual LLM reasoning
Researchers from Apple Inc. and the Hasso Plattner Institute have conducted a large-scale study on Group Relative Policy Optimization (GRPO) in non-English and multilingual settings. Their findings indicate that trainin…