PulseAugur
实时 21:47:45
English(EN) We evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested.

Perplexity AI:GPT-6 Astra 在 WANDR 基准测试中领先,超越 Fable 5.1 和 Opus 5

Perplexity AI 评估了 GPT-6 Astra,报告其在 WANDR 基准测试中的得分为 0.682,每项任务成本为 11.98 美元。这一性能代表了对其他测试模型的显著改进,比 Fable 5.1 的性能高出 13.5%,同时成本降低了 6.1%,并且比 Opus 5 的性能高出 27.0%,成本略高 3.3%。 AI

影响 为人工智能模型的性能和成本效益设定了新的基准,可能影响未来模型的开发和采用。

排序理由 该条目报告了在特定任务(WANDR)上对特定人工智能模型(GPT-6 Astra)的基准评估,并将其性能和成本与其他模型进行了比较。[lever_c_demoted from research: ic=1 ai=1.0]

在 X — Perplexity 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Perplexity AI:GPT-6 Astra 在 WANDR 基准测试中领先,超越 Fable 5.1 和 Opus 5

本文如何被排名

Signal score
19 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目报告了在特定任务(WANDR)上对特定人工智能模型(GPT-6 Astra)的基准评估,并将其性能和成本与其他模型进行了比较。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    我们使用 WANDR 评估了 GPT-6 Astra。其得分 0.682,每任务 11.98 美元,是所有我们测试过的模型中得分最高的。

    We evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at 3.3% higher cost. https://t.co/SyYmD38qvq