Researchers have introduced RealPref, a new benchmark designed to evaluate how well large language models (LLMs) can follow user preferences over extended interactions. The benchmark includes synthetic user profiles, personalized preferences, and long-horizon interaction histories, with varying types of preference expression from explicit to implicit. Initial results show a significant drop in LLM performance as context length increases and preferences become more implicit, highlighting challenges in generalizing user understanding to unseen scenarios. AI
IMPACT This benchmark could drive development of more personalized and adaptive AI assistants by highlighting current limitations in long-term preference following.
RANK_REASON The item is an academic paper introducing a new benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- large-language models
- LLMs
- Qianyun Guo
- RealPref
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →