A new research paper identifies a vulnerability in how large language models (LLMs) are evaluated and aligned, termed Rubric-Induced Preference Drift (RIPD). This occurs when edits to natural-language rubrics, even those that pass benchmark validation, can cause systematic shifts in an LLM judge's preferences. Researchers demonstrated that this drift can be exploited through "preference attacks," where benchmark-compliant rubric edits steer LLM judgments away from a trusted reference, reducing accuracy by up to 27.9%. This induced bias can propagate through alignment pipelines, leading to persistent and systematic drift in the behavior of trained LLM policies. AI
IMPACT Highlights a systemic alignment risk in LLMs, potentially impacting the reliability of AI-generated content and decisions.
RANK_REASON Research paper detailing a novel vulnerability in LLM evaluation methods. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- large-language models
- LLM judges
- Rubric-Induced Preference Drift
- Ruomeng Ding
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →