PulseAugur
EN
LIVE 22:53:03
ENTITY Group Relative Policy Optimisation

Group Relative Policy Optimisation

PulseAugur coverage of Group Relative Policy Optimisation — every cluster mentioning Group Relative Policy Optimisation across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
  1. TOOL · CL_228964 ·

    New LLM training and inference strategies boost Manim animation generation

    Researchers have developed a novel training pipeline called ManimTrainer, which combines supervised fine-tuning (SFT) with reinforcement learning (RL) techniques like Group Relative Policy Optimisation (GRPO). This appr…

  2. TOOL · CL_128824 ·

    New ASIG method enhances LLM information gathering via Bayesian design

    Researchers have developed a new fine-tuning approach called Amortised Sequential Information Gathering (ASIG) to improve how large language models (LLMs) gather information in sequential decision-making scenarios. ASIG…