A new study published on arXiv has found that preference judgments derived from large language models (LLMs) are not self-consistent. Researchers developed statistical tests to measure how well a single utility function could reproduce these judgments, finding significant inconsistencies across various examples like flight and apartment choices. This suggests that LLM-derived preference data cannot be reliably summarized by a single utility function, potentially impacting agent-based systems that rely on such data. AI
IMPACT Challenges the reliability of LLM-derived preference data for agent decision-making.
RANK_REASON Research paper published on arXiv detailing findings about LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Large Language Models
- Litmaps
- ScienceCast
- scite Smart Citations
- utility function
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →