PulseAugur
中
实时 23:28:37
English(EN) Evaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita

新的 Incognita 框架在社交任务中评估生成式代理

研究人员开发了 Incognita,一个用于在复杂社交任务环境中评估生成式代理的新框架。该系统基于康考迪亚大学,将社交互动与基于现实的执行分开,允许代理与中介动作的专家进行交流。Incognita-Retail 是一个特定应用,将 tau-Bench 零售环境改编为多实体设置。在 18 项任务上的评估表明,虽然代理在成功率方面有所提高并减少了过早定稿,但其整体可靠性仍然很低,突显了知识提取和动作理由方面的未来发展领域。 AI

影响 引入了一种新颖的评估方法,用于在复杂、社会分布式任务中评估生成式代理。

排序理由 学术论文,描述了一个新的生成式代理评估框架。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 Incognita 框架在社交任务中评估生成式代理

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,描述了一个新的生成式代理评估框架。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
93 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Dan C. Hsu, Luke Lu ·

    使用 Incognita 在社会分布式任务环境中评估基于动作的生成式代理

    arXiv:2607.02975v1 Announce Type: new Abstract: Effective agency in social environments depends on when an agent seeks knowledge, when it acts, and whether its actions are justified by acquired information. Existing grounded benchmarks provide executable actions, persistent state…