PulseAugur
实时 17:42:23

新的QUACK框架审计LLM代理的语言基础失败

研究人员推出了QUACK,这是一个开源环境和评估框架,旨在审计大型语言模型(LLM)代理在多模态社交推理任务中使用的语言基础。QUACK根据游戏结果、行为轨迹和话语级一致性来评估代理,特别标记空间幻觉和无根据指控等问题。对三个前沿VLM的评估显示出重大的基础失败,代理幻觉的空间声明超过15%,并且超过一半的指控没有证据支持。 AI

影响 该框架通过识别和纠正语言基础问题,可以提高LLM代理在复杂推理任务中的可靠性,从而使其更加健壮。

排序理由 该集群描述了一篇介绍LLM代理评估框架的新研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的QUACK框架审计LLM代理的语言基础失败

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一篇介绍LLM代理评估框架的新研究论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
103 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [3]

  1. arXiv cs.AI TIER_1 English(EN) · Ye Yuan, Rui Song, Weien Li, Zeyu Li, Haochen Liu, Xiangyu Kong, Changjiang Han, Yonghan Yang, Zichen Zhao, Zixuan Dong, Fuyuan Lyu, Bowei He, Haolun Wu, Jikun Kang, Xue Liu ·

    QUACK:多模态社交推理智能体中已沟通知识的提问、理解与审计

    arXiv:2605.27068v1 Announce Type: cross Abstract: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win…

  2. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xue Liu ·

    QUACK:多模态社交推理智能体中已沟通知识的提问、理解与审计

    Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win rates and largely remain to text-only interaction…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    QUACK:多模态社交推理智能体中已沟通知识的提问、理解与审计

    A multimodal social reasoning environment and evaluation framework called QUACK is introduced to audit the grounding of agent language through three-level assessment of game outcomes, behavioral trajectories, and utterance-level consistency.