A new research paper titled "Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates" has been published on arXiv. The study found that while Large Language Models (LLMs) can accurately predict average human responses to survey items, they fail to capture individual-specific variations. Even with richer persona data and fine-tuning, LLMs could only explain a small fraction of the respondent-specific variance, significantly underperforming human test-retest reliability. The research highlights that current LLMs exhibit "item-mean surrogacy," meaning they approximate item averages but not the nuanced, individual deviations required to truly substitute for humans. AI
IMPACT LLMs can approximate average human responses but cannot yet capture individual-specific variations, limiting their use as human surrogates.
RANK_REASON Research paper published on arXiv detailing findings about LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →