发表机构
Independent Researcher Seattle, WA, USA
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将LLM与开源形式化后端结合的多智能体RTL修复流水线,通过反例引导迭代实现RTL修复,经ALU案例及6基准测试验证可行性并分析失效模式。
AI 中文摘要
验证工作占据了现代芯片设计的大部分精力,但提供正确性数学保证的形式化验证工具仍价格高昂且许可限制严格。尽管大语言模型(LLM)已展现出硬件设计的应用潜力,现有RTL修复方法要么通过仅覆盖部分输入的仿真验证结果,要么依赖商业工具,且极少将形式化证明与完全开源的工具链结合。本文提出一种多智能体流水线,将LLM与开源形式化后端(Yosys、SymbiYosys和Z3)耦合,通过反例引导迭代修复RTL:该框架生成形式化属性、验证设计,并将反例反馈给LLM,直至设计通过k-归纳法证明正确或迭代预算耗尽。通过ALU案例研究,该流水线可检测并修复存在功能缺陷的设计并完成正确性形式化证明。在包含6个基准测试的套件中,1个设计被可靠修复,同时本文分析了4种不同失效模式:有界覆盖空值、规范歧义、时序逻辑缺陷及多属性压力。本文将此项工作视为带有详细失效分析的可行性研究,还报告了与开源形式化验证社区相关的Yosys bind指令的实际局限性。
英文摘要
Verification consumes the majority of modern chip design effort, yet the formal verification tools that provide mathematical guarantees of correctness remain expensive and restrictively licensed. While large language models (LLMs) have shown promise for hardware design, existing approaches to RTL repair validate their results through simulation - which exercises only a subset of inputs - or rely on commercial tools, and few combine formal proof with an entirely open-source toolchain. In this paper, we present a multi-agent pipeline that couples an LLM with an open-source formal backend (Yosys, SymbiYosys, and Z3) to repair RTL through counterexample-guided iteration: the framework generates formal properties, verifies the design, and feeds counterexamples back to the LLM until the design is proved correct by k-induction or an iteration budget is exhausted. Through an ALU case study, we show that the pipeline can detect and repair a real functional bug with a formal proof of correctness. Across a six-benchmark suite, one design is repaired reliably, and we characterize four distinct failure modes: bounded-cover vacuity, specification ambiguity, temporal-logic bugs, and multi-property pressure. We frame this work as a feasibility study with a detailed failure analysis, and additionally report a practical limitation of the Yosys bind directive relevant to the open-source formal verification community.
Comments6 pages, 3 figures