发表机构
Institute of Computing Technology, Chinese Academy of Sciences; University of the Chinese Academy of Sciences; School of Advanced Interdisciplinary Sciences(中国科学院计算技术研究所; 中国科学院大学; 先进交叉学科研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究硬件代理跨阶段设计封闭性评估难题,提出CLOSER-Bench协议,通过配对任务记录调用,测量多方面指标。经试点和验证揭示差距,构建加速器,促使将硬件封闭视为预算顺序决策问题。
AI 中文摘要
硬件工程使编码代理面临一种难以用单次通过率衡量的长期工作:进展是连续的,工具反馈延迟且异构,后端故障可能需要修改RTL而非调整其他物理设计参数。现有基准测试测量RTL生成、存储库修复、验证、PPA演变或物理实现,但不同设计和预言机难以确定代理在抽象边界处的成败。我们引入CLOSER-Bench,一种用于预算跨阶段设计封闭性的受控评估协议。针对一种设计和一个隐藏目标,它将规范到RTL、RTL到GDS以及规范到GDS任务配对,记录每个模拟器、综合、STA和布局布线调用,并测量最终质量、随时进展、工具成本和跨阶段恢复。该基准基于开源的Verilator、Yosys、OpenROAD.KLayout、Sky130和Harbor代理工具包构建。一个包含RTL修复、基于变异的验证、覆盖率、PPA优化、设计空间探索、跨模型调试和安全性的十任务试点建立了可执行工具包,并揭示了明显的完成-封闭差距:三个代理解决了局部AXI修复任务,而匹配的验证封闭任务将一个前沿代理与另外两个成功的基线区分开来。我们进一步验证了完整的RTL到GDS流程,并构建了一个基于宏的AXI/DMA流加速器用于阶段配对评估。这些结果促使将硬件封闭视为一个预算顺序决策问题,而非独立代码生成任务的集合。
英文摘要
Hardware engineering exposes coding agents to a form of long-horizon work that is difficult to capture with pass-at-k: progress is continuous, tool feedback is delayed and heterogeneous, and a backend failure may require revising RTL rather than tuning another physical-design parameter. Existing benchmarks measure RTL generation, repository repair, verification, PPA evolution, or physical implementation, but their different designs and oracles make it hard to determine where an agent succeeds or fails across abstraction boundaries. We introduce CLOSER-Bench, a controlled evaluation protocol for budgeted cross-stage design closure. For one design and one hidden objective, it pairs spec-to-RTL, RTL-to-GDS, and spec-to-GDS tasks, records every simulator, synthesis, STA, and place-and-route invocation, and measures final quality, anytime progress, tool cost, and cross-stage recovery. The benchmark is built on open-source Verilator, Yosys, OpenROAD, KLayout, Sky130, and the Harbor agent harness. A ten-task pilot spanning RTL repair, mutation-based verification, coverage, PPA optimization, design-space exploration, cross-model debugging, and security establishes the executable harness and exposes a sharp completion--closure gap: three agents solve a localized AXI repair task, while the matched verification-closure task separates a frontier agent from two otherwise successful baselines. We further validate a full RTL-to-GDS flow and construct a macro-based AXI/DMA streaming accelerator for the stage-paired evaluation. These results motivate treating hardware closure as a budgeted sequential decision problem rather than a collection of independent code generation tasks.
Comments6 pages, 2 figures, 4 tables