A new benchmark called RESCUE-BENCH has been introduced to evaluate large language models' ability to understand and respond to interpersonal dynamics in multi-party emotional support conversations. Constructed from real family and couple interviews, RESCUE-BENCH includes over 7,000 annotated turns and aims to assess relational understanding and relation-sensitive support. Experiments with ten LLMs revealed that while current models can handle basic emotional cues, they struggle with complex relational reasoning, indicating a gap in their capacity for nuanced interpersonal support. AI
IMPACT Highlights limitations in current LLMs for understanding and responding to complex interpersonal dynamics in conversations.
RANK_REASON The item describes a new benchmark for evaluating LLM capabilities in a specific research area. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →