ENTITY
AlpacaEval 2.0
AlpacaEval 2.0
PulseAugur coverage of AlpacaEval 2.0 — every cluster mentioning AlpacaEval 2.0 across labs, papers, and developer communities, ranked by signal.
Total · 30d
0
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
-
New method leverages reward model states for better AI feedback
Researchers have developed a new method called Representation-Aware Advantage Estimation (GraphAE) that enhances reinforcement learning from human feedback (RLHF). This technique utilizes the richer information encoded …
-
New S-SPPO framework enhances LLM alignment with human preferences
Researchers have introduced S-SPPO, a new framework designed to improve the alignment of large language models with human preferences. This method addresses instabilities in previous Self-Play Preference Optimization te…