Researchers have developed a new benchmark to evaluate Large Language Models' (LLMs) ability to understand social pragmatic inference in Chinese online comments. This benchmark, constructed from over 200,000 social media interactions, features 4,735 human-validated items designed to test if models can distinguish plausible readings of comments within their conversational context. The task proved challenging, with the best-performing model achieving 81.42% accuracy in a leave-writer-out setting, significantly lower than the 90.8% human accuracy. Analysis revealed that while models could often detect general irony or playfulness, they struggled to identify the specific interactional moves or mechanisms behind these nuances. AI
IMPACT This benchmark could drive improvements in LLMs' ability to understand subtle social cues and context in natural language, enhancing their utility in real-world communication scenarios.
RANK_REASON The item describes a new benchmark for evaluating LLMs on a specific NLP task, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →