PulseAugur
实时 22:21:27
English(EN) GPT-6 Astra is the new Extended NYT Connections Benchmark champion. GPT-6 Astra (xhigh) sets a new high score of 98.1, while high scores 97.7. Both outperform GPT-5.6 Sol while costing ~40% less per puzzle

GPT-6 Astra 在纽约时报连接基准测试中名列前茅,以更低的成本超越 GPT-5.6 Sol

GPT-6 Astra 在扩展的纽约时报连接基准测试中取得了 98.1 的新最高分,超越了之前的模型。一个变体 GPT-6 Astra (xhigh) 也获得了 97.7 的极高分数。与 GPT-5.6 Sol 相比,两个 GPT-6 Astra 模型都表现出卓越的性能,同时每个谜题的运行成本降低了约 40%。 AI

影响 在流行的基准测试中设定了新的性能标准,可能影响未来的模型开发和成本效益目标。

排序理由 模型的新基准分数。 [lever_c_demoted from research: ic=1 ai=1.0]

在 r/singularity 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

GPT-6 Astra 在纽约时报连接基准测试中名列前茅,以更低的成本超越 GPT-5.6 Sol

本文如何被排名

Signal score
8 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
模型的新基准分数。 [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/singularity TIER_2 English(EN) · /u/zero0_one1 ·

    GPT-6 Astra 成为新的纽约时报连接基准测试冠军。GPT-6 Astra (xhigh) 创下 98.1 的新高分,而高分则为 97.7。两者均优于 GPT-5.6 Sol,而每局成本却降低了约 40%

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1w7feo5/gpt6_astra_is_the_new_extended_nyt_connections/"> <img alt="GPT-6 Astra is the new Extended NYT Connections Benchmark champion. GPT-6 Astra (xhigh) sets a new high score of 98.1, while high scores 97.…