Researchers have introduced MeetingToM, a new benchmark designed to evaluate the theory of mind capabilities of multimodal large language models (MLLMs) in the context of multi-party meetings. This benchmark addresses limitations in existing evaluations by focusing on latent social states and group dynamics, such as pseudo-consensus where apparent agreement hides private dissent. MeetingToM assesses MLLMs at subject, dyadic, and group levels, with initial analyses showing persistent challenges for current models in integrating non-verbal cues and inferring hidden attitudes. AI
IMPACT This benchmark could drive advancements in MLLMs' ability to understand and navigate complex social interactions, crucial for applications in collaborative environments.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- MeetingToM
- Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
- theory of mind
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →