PulseAugur
实时 21:51:40
English(EN) CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents

新基准 CLOSER-Bench 评估硬件设计收敛中的 AI 代理

研究人员推出了 CLOSER-Bench,这是一个新的评估协议,旨在评估 AI 代理在硬件工程任务中的能力。该基准侧重于预算跨阶段设计收敛,整合了从规范到 RTL 生成,以及从 RTL 到物理实现 (GDS) 的任务。它利用了 VerilatorYosysOpenROAD 等开源工具,并衡量最终质量、随时间推移的进展、工具成本以及从后端故障中恢复的能力。初步测试显示,代理解决局部编码任务的能力与它们在集成验证和收敛挑战中的表现之间存在显著差距。 AI

影响 该基准可以通过提供标准化的评估框架,推动 AI 代理在复杂、长周期的工程任务方面的进步。

排序理由 该项目描述了一篇介绍用于评估特定领域(硬件工程)AI 代理基准的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准 CLOSER-Bench 评估硬件设计收敛中的 AI 代理

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Peilong Zhou, Zhirong Chen, Cangyuan Li, Haoyu Gao, Kaiyan Chang, Ziming Qu, Ying Wang ·

    CLOSER-Bench:评估硬件代理的预算跨阶段设计闭合

    arXiv:2607.16632v1 Announce Type: cross Abstract: Hardware engineering exposes coding agents to a form of long-horizon work that is difficult to capture with pass-at-k: progress is continuous, tool feedback is delayed and heterogeneous, and a backend failure may require revising …