A new report details the performance of the Qiushi Engine autonomous agent on the AstaBench E2E-Bench-Hard benchmark, utilizing DeepSeek's deepseek-v4pro-preview model. The engine achieved a score of 0.816 on the benchmark, with a cost of $15.209 per task. While it demonstrated a 10% full-task completion rate, surpassing previous agent performance by 7 percentage points, it satisfied 82.1% of the required rubric items across 40 tasks. The evaluation highlighted limitations in areas such as repeated runs, external dependencies, and ablation studies, despite successfully producing and verifying reports, code, and experimental artifacts. AI
IMPACT Demonstrates progress in autonomous agent capabilities for complex research tasks, highlighting areas for future development.
RANK_REASON Research paper detailing performance on an AI benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →