PulseAugur
EN
LIVE 19:17:53
ENTITY Reinforcement Learning with Verifiable Rewards

Reinforcement Learning with Verifiable Rewards

PulseAugur coverage of Reinforcement Learning with Verifiable Rewards — every cluster mentioning Reinforcement Learning with Verifiable Rewards across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
17
44 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
17
43 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

11 day(s) with sentiment data

RECENT · PAGE 1/3 · 44 TOTAL
  1. TOOL · CL_196030 ·

    CuSearch framework enhances agentic RAG training with curriculum sampling

    Researchers have developed CuSearch, a new framework for training agentic retrieval-augmented generation (RAG) systems using Reinforcement Learning with Verifiable Rewards (RLVR). This method addresses the issue of unif…

  2. RESEARCH · CL_193290 ·

    New research tackles LLM reasoning reliability and hallucination

    Multiple research papers explore methods to enhance the reliability and accuracy of Large Language Models (LLMs) in reasoning tasks. One approach, REIN, uses reflection and abstention to reduce hallucinations by allowin…

  3. RESEARCH · CL_193355 ·

    New RLVR methods enhance LLM robustness and generalization · 2 sources tracked

    Researchers have developed new methods to improve the robustness and generalization of Reinforcement Learning with Verifiable Rewards (RLVR) for Multimodal Large Language Models. The first approach, Prompt-Invariant RLV…

  4. TOOL · CL_193343 ·

    New framework enhances medical mathematical reasoning with knowledge guidance

    Researchers have introduced MedCalc-R1, a novel knowledge-guided reward framework designed to improve mathematical reasoning in medical contexts. This framework addresses limitations in existing Reinforcement Learning w…

  5. RESEARCH · CL_193326 ·

    New benchmarks and methods enhance multimodal AI reasoning and trustworthiness · 4 sources tracked

    Researchers are developing new methods to improve the reliability and trustworthiness of multimodal large language models (MLLMs). One approach, VERDICT, uses disagreement among multiple verifiers to identify errors in …

  6. TOOL · CL_178378 ·

    New GeoRA method enhances RLVR for large language models

    Researchers have introduced GeoRA, a novel low-rank adaptation method specifically designed for Reinforcement Learning with Verifiable Rewards (RLVR). Unlike existing methods that focus on supervised fine-tuning, GeoRA …

  7. RESEARCH · CL_180478 ·

    New method RSTG improves LLM reinforcement learning with adaptive teacher guidance

    Researchers have developed RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), a novel method to improve reinforcement learning for large language models. Existing methods like GRPO struggle with sparse re…

  8. TOOL · CL_174168 ·

    New reward functions boost AI model unlearning efficiency for privacy compliance

    Researchers have developed new reward functions for machine unlearning, a process that selectively removes specific knowledge from AI models. This is crucial for complying with privacy regulations like GDPR and the EU A…

  9. TOOL · CL_179300 ·

    New method CSCR improves LLM long-context reasoning by reallocating token credit

    Researchers have developed a new method called Counterfactual Sensitivity Credit Reallocation (CSCR) to improve the reasoning capabilities of large language models, particularly in tasks requiring long-context reasoning…

  10. TOOL · CL_156396 ·

    Research paper analyzes reasoning's impact on LLM translation quality

    A new research paper explores the effectiveness of Reinforcement Learning with Verifiable Rewards (RLVR) for training large-language models (LLMs), particularly for Neural Machine Translation (NMT). The study investigat…

  11. TOOL · CL_154399 ·

    New RLVR Framework PACS Enhances LLM Reasoning Capabilities

    Researchers have introduced PACS, a novel framework for Reinforcement Learning with Verifiable Rewards (RLVR) designed to improve the reasoning capabilities of large language models (LLMs). PACS reformulates RLVR as a s…

  12. RESEARCH · CL_156464 ·

    New H$^2$SD framework boosts LLM reasoning via hybrid self-distillation

    Researchers have developed H$^2$SD, a novel hybrid hindsight self-distillation framework designed to enhance the reasoning abilities of large language models. This method addresses limitations in existing reinforcement …

  13. RESEARCH · CL_154008 ·

    New research explores reinforcement learning advancements across multiple domains · 10 sources tracked

    Multiple research papers published on arXiv explore advancements in reinforcement learning (RL) and its applications. One study focuses on improving the interpretability of RL policies through decision-tree pruning, dem…

  14. RESEARCH · CL_147463 ·

    New Contrastive Policy Optimization method improves reinforcement learning

    Researchers have introduced Contrastive Policy Optimization (CPO), a novel method for reinforcement learning with verifiable rewards. CPO utilizes token-level contrastive disagreement between generated text distribution…

  15. TOOL · CL_152483 ·

    New Contrastive Policy Optimization framework enhances reinforcement learning

    Researchers have introduced Contrastive Policy Optimization (CPO), a novel framework for reinforcement learning with verifiable rewards. CPO leverages token-level contrastive disagreement between reference-guided and va…

  16. RESEARCH · CL_145746 ·

    SIVA-RL framework enhances multimodal reasoning by grounding predictions in visual evidence

    Researchers have introduced SIVA-RL, a novel framework designed to improve multimodal reinforcement learning by ensuring vision-language models ground their predictions in visual evidence. Unlike previous methods that r…

  17. TOOL · CL_129050 ·

    New AI training method counts errors instead of rubrics for subjective tasks

    Researchers have introduced Implicit Error Counting (IEC), a novel method for training AI models in tasks where ideal outputs are subjective or non-existent. Unlike traditional reward systems that focus on correctness a…

  18. TOOL · CL_128681 ·

    LLMs learn evidence-seeking diagnostic reasoning with reinforcement learning

    Researchers have developed a new framework that uses Reinforcement Learning with Verifiable Rewards (RLVR) to enable Large Language Models (LLMs) to perform evidence-seeking diagnostic reasoning. This approach addresses…

  19. RESEARCH · CL_128417 ·

    New research explores controllable generalization failures and efficient RL distillation for LLMs

    Researchers are exploring new methods to improve language model generalization and reasoning capabilities. One paper proposes a technique to construct models that exhibit controllable generalization failures by training…

  20. RESEARCH · CL_123196 ·

    New framework enhances LLM evaluation with multi-role rubric generation

    Researchers have introduced Multi-Role Rubric Generation (MRRG), a novel framework designed to improve the evaluation of large language models (LLMs) on open-ended tasks. Unlike previous methods that rely on a single ev…