PulseAugur
实时 14:02:55
English(EN) MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

新的基准MeetingToM测试多模态大语言模型在社交推理方面的能力

研究人员推出了MeetingToM,这是一个旨在评估多模态大语言模型(MLLMs)在多人会议情境下心智理论能力的新基准。该基准通过关注潜在的社交状态和群体动态(例如,表面上一致但私下有异议的伪共识)来解决现有评估的局限性。MeetingToM在个体、二元和群体层面评估MLLMs,初步分析显示当前模型在整合非语言线索和推断隐藏态度方面仍面临持续挑战。 AI

影响 该基准有望推动MLLM理解和驾驭复杂社交互动能力的发展,这对于协作环境中的应用至关重要。

排序理由 该集群描述了一个用于评估AI模型的新学术基准。[lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的基准MeetingToM测试多模态大语言模型在社交推理方面的能力

报道来源 [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

    Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across sp…