arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17528cs.AIcs.ARcs.LG

人工智能代理真的能完成从RTL到GDS的转换吗?基准测试工具交互式EDA工作流程的经验教训

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows

Jinyuan Deng, Zhengrui Chen, Xufeng Wei, Tianyu Xing, Chenyi Wen, Qi Sun, Cheng Zhuo

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨人工智能代理在EDA工作流程中的表现,提出FluxBench进行系统评估,涵盖多种场景并评估多项能力,引入Token ROI指标。结果显示不同代理系统架构性能有差距,特定领域技能不是提升性能的唯一关键,系统设计和基础模型能力也很重要。

中文摘要 AI 辅助

基于大语言模型的代理系统已成为电子设计自动化(EDA)的一种有前景的范式,在自动化复杂设计工作流程方面展现出强大潜力。但现有评估主要针对孤立的EDA任务考察单个语言模型,对不同代理系统在完整EDA流程中的表现了解有限。本文提出FluxBench,在统一提示、工具环境和技术库设置下,对人工智能代理进行端到端EDA工作流程的系统评估。评估涵盖代表性场景,包括用开源工具链生成RTL以及使用闭源商业EDA工具进行工业应用的从RTL到GDS流程。通过这些工作流程,评估代理在RTL代码生成、迭代修复、工具反馈利用、逻辑综合、布局布线和工程变更单自动化等方面的能力。为进一步刻画代理系统的效率,引入令牌投资回报率(Token ROI)这一成本效率指标,衡量相对于令牌使用和运行时成本,EDA工件的有效改进。实验结果表明,即使基于相同基础模型,不同代理系统架构的性能差距可达86.27%。在具有可比任务性能的系统中,Token ROI的差异可达105.92倍。以使用PicoRV32的从RTL到GDS流程为例,FluxEDA的端到端得分高达97.94,比具备特定领域EDA技能的Claude Code高出8.39倍。这些结果表明,仅特定领域技能不足以在大规模EDA场景中提高代理性能。相反,代理系统设计和基础模型能力在实现有效的自动化EDA工作流程中都起着关键作用。

英文摘要

Large language model (LLM) agents are extending electronic design automation (EDA) beyond static RTL generation toward long-horizon, tool-interactive workflows. Yet it remains unclear whether general-purpose coding agents, even with domain-specific EDA skills, can reliably execute an end-to-end RTL-to-GDS flow encompassing synthesis, physical implementation, and engineering change order (ECO) optimization. We evaluate AI agents on a PicoRV32 RTL-to-GDS flow using commercial EDA tools under two timing targets. Their performance is assessed using end-to-end design score, stage completion, and Token ROI, a cost-efficiency metric relating design quality to runtime and cost. Comparing three agent architectures and four foundation models, we derive three practical lessons. First, domain-specific skills improve agents' understanding of individual subtasks but do not ensure reliable completion of a long-horizon EDA flow. Second, agents that achieve similar design progress can still differ by up to 141 times in Token ROI, revealing substantial differences in runtime and cost efficiency. Third, low-level tool-interface mismatches are a major source of physical design failures, particularly when Tcl commands depend on the tool version or execution mode. These results suggest that robust Agentic EDA requires not only stronger models but also structured tool interfaces, persistent design context, controlled execution, and process-level evaluation.

↑