ENTITY
Group Relative Policy Optimisation
Group Relative Policy Optimisation
PulseAugur coverage of Group Relative Policy Optimisation — every cluster mentioning Group Relative Policy Optimisation across labs, papers, and developer communities, ranked by signal.
Total · 30d
2
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
-
New LLM training and inference strategies boost Manim animation generation
Researchers have developed a novel training pipeline called ManimTrainer, which combines supervised fine-tuning (SFT) with reinforcement learning (RL) techniques like Group Relative Policy Optimisation (GRPO). This appr…
-
New ASIG method enhances LLM information gathering via Bayesian design
Researchers have developed a new fine-tuning approach called Amortised Sequential Information Gathering (ASIG) to improve how large language models (LLMs) gather information in sequential decision-making scenarios. ASIG…