A new arXiv paper compares Evolution Strategies (ES) and Group Relative Policy Optimization (GRPO) for post-training large language models. While both methods achieve comparable accuracy on single-task and sequential learning scenarios, they produce significantly different model updates. ES induces larger, broader changes with more off-task KL drift, whereas GRPO makes smaller, more localized adjustments. The research suggests that gradient-free and gradient-based fine-tuning can arrive at similarly accurate but geometrically distinct solutions, impacting knowledge preservation and forgetting. AI
IMPACT This research highlights how different fine-tuning methods can lead to distinct model behaviors, potentially impacting knowledge retention and transfer in LLMs.
RANK_REASON The cluster contains an academic paper detailing a comparison of two methods for LLM post-training. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →