PulseAugur
EN
LIVE 08:58:04

Qiushi Engine achieves 10% completion rate on AstaBench benchmark

A new report details the performance of the Qiushi Engine autonomous agent on the AstaBench E2E-Bench-Hard benchmark, utilizing DeepSeek's deepseek-v4pro-preview model. The engine achieved a score of 0.816 on the benchmark, with a cost of $15.209 per task. While it demonstrated a 10% full-task completion rate, surpassing previous agent performance by 7 percentage points, it satisfied 82.1% of the required rubric items across 40 tasks. The evaluation highlighted limitations in areas such as repeated runs, external dependencies, and ablation studies, despite successfully producing and verifying reports, code, and experimental artifacts. AI

IMPACT Demonstrates progress in autonomous agent capabilities for complex research tasks, highlighting areas for future development.

RANK_REASON Research paper detailing performance on an AI benchmark. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qiushi Engine achieves 10% completion rate on AstaBench benchmark

How we ranked this

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Research paper detailing performance on an AI benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Wenhao Li, Shuxing Yang, Fujia Chen, Jincheng Mi, Yuang Pan, Rui Zhao, Zichen Li, Junyao Wu, Shenzhan Hong, Yaqi Li, Yize Wang, Kaihao Zhu, Taowen Deng, Junjie Yang, Hongsheng Chen, Yihao Yang ·

    Qiushi Engine on AstaBench E2E-Bench-Hard

    arXiv:2609.08196v1 Announce Type: new Abstract: This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark that requires autonomous agents to carry a research question through experimental design, code implementation, actual executio…