PulseAugur
实时 00:54:12
中文(ZH) GPT-5.6-sol入榜DRACO:OpenSquilla集成方案仍在Brave组质量、成本双领先

OpenSquilla 集成在成本和质量上优于 GPT-5.6-sol 在 DRACO 基准测试上

DRACO 基准测试的 Brave Search 组的最新评估显示,GPT-5.6-sol 已进入排行榜,平均得分为 63.99,每任务成本为 1.71 美元。然而,一个名为 OpenSquilla 0.5.0 Preview 的集成解决方案,利用了多个国内中文模型,以显著更低的每任务 0.12 美元成本取得了略高的 64.09 分。这表明复杂代理任务的竞争正从依赖单一强大模型转向有效组织多个模型。 AI

影响 强调了模型编排在复杂代理任务中日益增长的重要性,超越了单一模型的优势。

排序理由 该项目报告了 AI 模型的基准测试结果,特别比较了在 DRACO 基准测试上的性能和成本。[lever_c_demoted from research: ic=1 ai=1.0]

在 量子位 (QbitAI) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

OpenSquilla 集成在成本和质量上优于 GPT-5.6-sol 在 DRACO 基准测试上

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目报告了 AI 模型的基准测试结果,特别比较了在 DRACO 基准测试上的性能和成本。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
57 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    GPT-5.6-sol 进入 DRACO 排行榜:OpenSquilla 集成解决方案仍在质量和成本方面领先 Brave 团队