Researchers have introduced MeetingToM, a new benchmark designed to evaluate the Theory of Mind (ToM) capabilities of Multimodal Large Language Models (MLLMs) in the context of multi-party meetings. This benchmark addresses limitations in existing evaluations by focusing on latent social states and group dynamics, such as pseudo-consensus, where apparent agreement masks private dissent. MeetingToM assesses ToM at subject, dyadic, and group levels, with analyses of current MLLMs revealing persistent challenges in integrating non-verbal cues and inferring hidden attitudes. AI
IMPACT This benchmark could drive advancements in AI's ability to understand and participate in complex social interactions.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Hugging Face
- MeetingToM
- Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
- pseudo-consensus
- theory of mind
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →