Researchers have introduced LUNAR, a novel benchmark designed to evaluate how well large language models can personalize responses based on longitudinal user behavior logs. Unlike previous benchmarks that relied on simple personas or isolated signals, LUNAR incorporates heterogeneous daily-life activities across domains such as clothing, food, housing, and mobility. The benchmark is constructed using a multi-stage synthesis pipeline to address data sparsity and privacy concerns, showing closer alignment with real behavioral distributions than other synthetic benchmarks. Experiments with 19 mainstream LLMs revealed that while access to behavioral logs is crucial, it is not sufficient for deep personalization, with evidence selection and cross-domain integration emerging as key challenges. AI
IMPACT This benchmark could drive improvements in how LLMs understand and respond to individual user behaviors, leading to more tailored and useful AI applications.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Large Language Models
- Litmaps
- LUNAR
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →