PulseAugur
实时 02:51:15
English(EN) Schema: a harness for llms, with Fable+4.8 or GPT 5.6 Sol, (supposedly) achieves 99% and 95.35% respectively on ARC-AGI-3.

Schema 框架声称在 ARC-AGI 3 基准测试中取得高分

一个名为 Schema 的新框架已被推出,旨在评估大型语言模型。早期报告表明,Schema 在使用 Fable+4.8 和 GPT 5.6 Sol 等模型时,在 ARC-AGI 3 基准测试中分别取得了 99% 和 95.35% 的令人印象深刻的分数。然而,一项澄清表明这些分数是在公共数据集上取得的,在保留集上的表现仍有待观察。 AI

影响 该框架可以为评估 LLM 在复杂推理任务上的能力提供新的标准。

排序理由 该项目描述了一个用于评估 LLM 的新框架并报告了基准分数,这属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/singularity 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Schema 框架声称在 ARC-AGI 3 基准测试中取得高分

报道来源 [1]

  1. r/singularity TIER_2 English(EN) · /u/TFenrir ·

    Schema:一个用于大型语言模型的框架,据称使用 Fable+4.8 或 GPT 5.6 Sol 分别在 ARC-AGI-3 上达到了 99% 和 95.35% 的准确率。

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1uyd4g9/schema_a_harness_for_llms_with_fable48_or_gpt_56/"> <img alt="Schema: a harness for llms, with Fable+4.8 or GPT 5.6 Sol, (supposedly) achieves 99% and 95.35% respectively on ARC-AGI-3." src="https://p…