A new study published on arXiv suggests that large language models (LLMs) exhibit a higher degree of agreement among themselves than they do with human readers. Researchers measured the convergence of LLMs against a reference set of human annotations, finding that models consistently agreed more with each other than with how humans naturally highlight text. This phenomenon was observed across various models from different vendors and countries, with even frontier models showing twice the agreement with themselves compared to GPT-4o's self-agreement. AI
IMPACT Suggests LLMs may be developing internal consensus that diverges from human interpretation, potentially impacting their utility in tasks requiring nuanced human alignment.
RANK_REASON Research paper published on arXiv detailing findings about language model behavior.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →