PulseAugur
中
实时 00:27:23
English(EN) MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

新的基准测试 MeetingToM 评估大语言模型在会议中的社交推理能力

研究人员推出了 MeetingToM,这是一个新的基准测试,旨在评估多模态大语言模型(MLLMs)在多方会议情境下的心智理论(ToM)能力。该基准测试通过关注潜在的社交状态和群体动态(例如,表面上的一致但私下有异议的伪共识)来解决现有评估的局限性。MeetingToM 在个体、二元和群体层面评估 ToM,对当前 MLLMs 的分析显示,在整合非语言线索和推断隐藏态度方面仍然存在挑战。 AI

影响 该基准测试有望推动 AI 在理解和参与复杂社交互动方面的能力取得进步。

排序理由 该集群描述了一篇介绍用于评估 AI 模型基准测试的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的基准测试 MeetingToM 评估大语言模型在会议中的社交推理能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇介绍用于评估 AI 模型基准测试的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
78 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Ziyi Wang, Yuhang Wu, Dongxu Piao, Xingyu Liu, Tianhui Zhou, Miao Liu ·

    MeetingToM:在多方会议中评估多模态大语言模型在心智理论推理上的表现

    arXiv:2607.19235v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-par…