A new evaluation framework called CoCoEval has been developed to assess large language models (LLMs) in simulating human social interactions, specifically focusing on inconsistent and uncollaborative behaviors. Researchers found that current LLMs like GPT-4.1, GPT-5.1, and Claude Opus 4 exhibit significantly fewer such behaviors than humans, and that prompt engineering and fine-tuning do not reliably bridge this gap. The study raises concerns about the accuracy of LLMs as proxies for genuine human social dynamics. AI
IMPACT Highlights limitations in LLM's ability to accurately simulate nuanced human social interactions, impacting their use in social science research.
RANK_REASON Research paper introducing a new evaluation framework for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →