A Reddit user is attempting to reproduce OpenAI's "persistently beneficial models" research but is encountering difficulties installing a desired trait using GRPO. The user's GRPO training run only achieved a minor +2.4 point increase in the trait, falling short of the ~+15 points needed. They have ruled out common issues like reward hacking, memorization, and dead gradients, and a co-author suggested that the number of distinct trait prompts used (20) was likely too few. The user is seeking advice on prompt count, rubric grading, and whether stylistic traits install differently than task-oriented ones. AI
IMPACT Highlights challenges in reproducing advanced RLHF techniques for model alignment.
RANK_REASON User attempting to reproduce a published research paper's findings and seeking community help. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →