arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大型语言模型(LLM)的硬件开发:分层中间表示(IR)与端到端多智能体工作流

LLM-based Hardware Development with Hierarchical IRs and End-to-End Multi-Agent Workflow

Chenyang Yin, Agasthi Haputhanthri, Aditya Anirudh Jonnalagadda, Zhenyu Bai, Yuanming Song, Saranyu Chattopadhyay, Mohammad Fadiheh, Tom Zelazny, Subhasish Mitra, Tulika Mitra

arXiv 2608.30659首次发表:更新:

发表机构

National University of Singapore; Shandong University; Stanford University; LUBIS EDA(新加坡国立大学; 山东大学; 斯坦福大学; LUBIS EDA)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM在硬件设计中应用受限的问题,本文提出基于分层IR与多智能体工作流的LLM硬件开发框架,在Verilog-Eval基准获95.5% pass@5,可生成符合标准的功能性复杂硬件设计。

AI 中文摘要

大型语言模型(LLM)在软件开发中的应用日益广泛,但其在复杂硬件设计中的应用仍有限。这一差距源于公开硬件训练数据的稀缺,以及硬件设计所采用的方法存在根本差异。特别是,将LLM应用于硬件需要的不仅是直接生成寄存器传输级(RTL)代码:模型必须理解模块边界、模块间连接以及验证要求。在本文中,我们提出了一种基于LLM的硬件开发框架,该框架采用分层中间表示(IR)与端到端多智能体工作流。核心思路是通过两种结构化IR为LLM提供硬件设计的抽象:架构草图(Architectural Sketch),用于捕获模块拓扑结构与互连;操作规范(Operational Specification),用于定义每个模块的功能与接口。我们的框架利用这些IR将复杂设计分解为子模块,指定每个模块的功能,并推导每个模块应如何测试与验证。我们在框架中引入了多智能体调试循环,使智能体能够获取错误反馈并控制调试细节,例如模拟时需探测的信号。我们在Verilog-Eval基准上对框架进行评估,获得了95.5%的pass@5通过率,超过了当前最先进的LLM生成框架。为更好评估框架在复杂、实际设计上的性能,我们引入了一项新的案例研究,涵盖从通用处理器到数字信号处理系统的应用。实验结果表明,这类复杂设计超出了现有方法的能力,而我们的框架是唯一能够生成功能性端到端设计的框架。我们生成的RTL符合所有行业标准设计规则,无 lint 错误,功能正确且完全可综合。

英文摘要

Large language models (LLMs) are increasingly used in software development, but their use in complex hardware design remains limited. This gap stems from both the scarcity of public hardware training data and the fundamentally different methodologies used in hardware design. In particular, applying LLMs to hardware requires more than direct RTL generation: the model must understand module boundaries, inter-module connections, and verification requirements. In this paper, we present an LLM-based hardware development framework with hierarchical intermediate representations (IRs) and an end-to-end multi-agent workflow. The core idea is to provide an abstraction of hardware design to LLMs through two structured IRs: Architectural Sketch, which captures module topology and interconnection, and Operational Specification, which defines per-module functionality and interfaces. Our framework uses these IRs to decompose a complex design into sub-modules, specify the per-block functionality, and derive how each module should be tested and verified. We incorporate a multi-agent debug loop in the framework, allowing agents to get the error feedback and control the debug details such as the signals to be probed for simulation. We evaluate our framework on Verilog-Eval benchmark, achieving a pass@5 rate of 95.5%, which surpasses current state-of-the-art LLM generation frameworks. To better assess performance on complex, realistic designs, we introduce a new case study spanning applications from general-purpose processors to digital signal processing systems. Experimental results indicate that such complex designs exceed the capabilities of existing approaches, whereas our framework is the only one capable of producing functional end-to-end design. Our generated RTL follows all industry-standard design rules, is lint-clean, functionally correct and fully synthesizable.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑