PulseAugur
实时 08:25:07
English(EN) Qiushi Engine on AstaBench E2E-Bench-Hard

求是引擎在 AstaBench 基准测试中完成率达到 10%

一份新报告详细介绍了求是引擎自主代理在 AstaBench E2E-Bench-Hard 基准测试上的表现,该测试使用了 DeepSeekdeepseek-v4pro-preview 模型。该引擎在该基准测试上得分 0.816,每任务成本为 15.209 美元。虽然其 10% 的完整任务完成率超过了之前代理性能 7 个百分点,但在 40 个任务中满足了 82.1% 的必需评分项。评估强调了在重复运行、外部依赖和消融研究等方面的局限性,尽管成功生成并验证了报告、代码和实验产物。 AI

影响 展示了自主代理在复杂研究任务方面的能力进展,并指出了未来发展的方向。

排序理由 关于在 AI 基准测试上表现的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

求是引擎在 AstaBench 基准测试中完成率达到 10%

本文如何被排名

Signal score
18 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
关于在 AI 基准测试上表现的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Wenhao Li, Shuxing Yang, Fujia Chen, Jincheng Mi, Yuang Pan, Rui Zhao, Zichen Li, Junyao Wu, Shenzhan Hong, Yaqi Li, Yize Wang, Kaihao Zhu, Taowen Deng, Junjie Yang, Hongsheng Chen, Yihao Yang ·

    Qiushi Engine 在 AstaBench E2E-Bench-Hard 上

    arXiv:2609.08196v1 Announce Type: new Abstract: This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark that requires autonomous agents to carry a research question through experimental design, code implementation, actual executio…