PulseAugur
中
实时 09:57:38
English(EN) ProgramBench result for Fable 5 is in, doubling Opus 4.8 even with 4.8 fallback "99% of the runs"

Fable 5 基准测试显示性能是 Opus 4.8 的两倍

一项新的基准测试 ProgramBench 已用于评估 Fable 5,结果表明其性能显著优于 Opus 4.8。基准测试的创建者指出,即使 Fable 5 在某些任务中使用了回退机制至 Opus 4.8,其性能仍是 Opus 4.8 的两倍。一个有趣的观察是,Fable 5 中回退到 Opus 4.8 所消耗的 token 是 Opus 4.8 单独执行类似任务的两倍。 AI

影响 Fable 5 在 ProgramBench 上性能达到 Opus 4.8 的两倍,表明其能力有了显著飞跃,可能给竞争对手带来压力。

排序理由 该集群报告了特定 AI 模型的基准测试结果,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/singularity 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Fable 5 基准测试显示性能是 Opus 4.8 的两倍

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群报告了特定 AI 模型的基准测试结果,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
113 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/singularity TIER_2 English(EN) · /u/reefine ·

    Fable 5 的 ProgramBench 结果已出,即使在 4.8 回退“99% 的运行”情况下也使 Opus 4.8 翻倍

    <!-- SC_OFF --><div class="md"><p><a href="https://x.com/ValsAI/status/2066760552156971291">https://x.com/ValsAI/status/2066760552156971291</a></p> <p>Quite interesting result, ProgramBench creator seem to imply that there is a difference between Fable 5 falling back to 4.8 quick…