A new study published on arXiv investigates the evaluative tendencies of large language models (LLMs) by examining their preferences for films. Researchers found that eight models from Anthropic, OpenAI, Alibaba Group, and Mistral AI consistently favored critically acclaimed but commercially obscure films over those that were commercially successful but critically unrecognized. This critical acclaim orientation was observed to increase with model scale within each family. The study also indicated that prompt framing significantly influences model rankings, suggesting that critical acclaim bias may manifest indirectly in real-world LLM applications. AI
IMPACT Suggests LLMs may exhibit biases similar to human critics, potentially influencing recommendation systems and content generation.
RANK_REASON Academic paper on LLM behavior and evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
- Alibaba Group
- Anthropic
- arXiv
- Bradley--Terry model
- Hugging Face
- Large Language Models
- Ordinary Least Squares
- Mistral AI
- OpenAI
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →