PulseAugur
中
实时 21:14:09
English(EN) Schema: a harness for llms, with Fable+4.8 or GPT 5.6 Sol, (supposedly) achieves 99% and 95.35% respectively on ARC-AGI-3.

Schema 框架声称在 ARC-AGI 3 基准测试中取得高分

一个名为 Schema 的新框架已被推出,旨在评估大型语言模型。早期报告表明,Schema 在使用 Fable+4.8 和 GPT 5.6 Sol 等模型时,在 ARC-AGI 3 基准测试中分别取得了 99% 和 95.35% 的令人印象深刻的分数。然而,一项澄清表明这些分数是在公共数据集上取得的,在保留集上的表现仍有待观察。 AI

影响 该框架可以为评估 LLM 在复杂推理任务上的能力提供新的标准。

排序理由 该项目描述了一个用于评估 LLM 的新框架并报告了基准分数,这属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/singularity 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Schema 框架声称在 ARC-AGI 3 基准测试中取得高分

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目描述了一个用于评估 LLM 的新框架并报告了基准分数,这属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
75 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/singularity TIER_2 English(EN) · /u/TFenrir ·

    Schema:一个用于大型语言模型的框架,据称使用 Fable+4.8 或 GPT 5.6 Sol 分别在 ARC-AGI-3 上达到了 99% 和 95.35% 的准确率。

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1uyd4g9/schema_a_harness_for_llms_with_fable48_or_gpt_56/"> <img alt="Schema: a harness for llms, with Fable+4.8 or GPT 5.6 Sol, (supposedly) achieves 99% and 95.35% respectively on ARC-AGI-3." src="https://p…