PulseAugur
实时 17:53:25
English(EN) Hi, I’m Troy. I build open-source tools in R. My agent harness, corteza, scored 98.6% on ARC-AGI-3’s 25 public games with Claude Opus 5 at xhigh, completing all

开源工具 Corteza 在 ARC-AGI-3 上以 Claude Opus 5 取得 98.6% 的分数

一个名为 Corteza 的开源工具,由 Troy 开发,在 ARC-AGI-3 基准测试中取得了 98.6% 的分数。这一成就的实现是使用了 Claude Opus 5,并完成了 25 个公开游戏中的全部 183 个关卡。Corteza 套件旨在在一个持久的 R 工作空间中编写和重用函数,允许数据和函数在对话压缩过程中持续存在。 AI

影响 展示了 LLM 在复杂基准测试中的高级推理能力,可能影响未来 AI 的发展。

排序理由 使用 LLM 进行的研究基准测试结果。[lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

开源工具 Corteza 在 ARC-AGI-3 上以 Claude Opus 5 取得 98.6% 的分数

本文如何被排名

Signal score
16 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
使用 LLM 进行的研究基准测试结果。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. Mastodon — mastodon.social TIER_1 English(EN) · troyhernandez ·

    你好,我是 Troy。我用 R 语言构建开源工具。我的代理套件 corteza,在使用 Claude Opus 5 以 xhigh 模式运行,并在 ARC-AGI-3 的 25 个公开游戏中取得了 98.6% 的分数,完成了所有任务

    Hi, I’m Troy. I build open-source tools in R. My agent harness, corteza, scored 98.6% on ARC-AGI-3’s 25 public games with Claude Opus 5 at xhigh, completing all 183 levels. The harness writes and reuses functions in a persistent R workspace. Functions and data survive conversatio…