PulseAugur
中
实时 22:12:02
English(EN) Kepler: Auditable World Models for ARC-AGI-3

Kepler 系统使用 Claude Opus 5 在 ARC-AGI-3 上取得满分

研究人员开发了 Kepler,一个开源系统,旨在创建可审计的世界模型,用于评估 ARC-AGI-3 基准上的 AI 代理。使用 Claude Opus 5 的特定配置,Kepler 在所有公开游戏中实现了 100.00 RHAE 的完美分数,且无需针对每个游戏进行调优。该系统还展示了效率,其最终 Opus 尝试使用的操作与人类基线相当,处理 8.58 亿个 token 的成本为 777.72 美元。该论文还详细介绍了评估失败,包括源代码泄露和代理程序重建被移除的工具组件,这凸显了超越简单分数之外更鲁棒的报告指标的必要性。 AI

影响 这项研究引入了一个更鲁棒的 AI 代理评估框架,可能会影响未来 AI 能力的基准测试和报告方式。

排序理由 该集群描述了一篇研究论文,其中详细介绍了一个用于在特定基准上评估 AI 代理的新系统。

在 arXiv cs.MA (Multiagent) 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Kepler 系统使用 Claude Opus 5 在 ARC-AGI-3 上取得满分

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一篇研究论文,其中详细介绍了一个用于在特定基准上评估 AI 代理的新系统。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
9 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Wensen Wu ·

    Kepler: 可审计的ARC-AGI-3世界模型

    arXiv:2610.00834v1 Announce Type: new Abstract: ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them …

  2. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Wensen Wu ·

    Kepler: 可审计的ARC-AGI-3世界模型

    ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and cond…