AI 中文总结
研究基于大语言模型的Verilog生成中不完美规范问题,提出VClare框架,通过规范级和仿真级两种修复范式修复缺陷,利用新基准数据集实验,单模块任务使设计通过率提高12.7%,多模块任务提高13.7%,提升了大语言模型在硬件设计中的应用潜力。
AI 中文摘要
大语言模型在根据自然语言规范生成Verilog代码方面展现出了有前景的能力。然而,人工编写的规范常包含语义缺陷,会显著降低大语言模型生成的硬件设计质量。本文首次对不完美规范进行系统研究并提出自动化框架VClare来修复它们,以提高生成的Verilog设计质量。该框架探索了两种互补的修复范式,即规范级修复和仿真级修复。此外,还提出了两个带有系统注入规范缺陷的新基准数据集。实验表明,对于单模块任务,VClare框架能有效修复规范缺陷,使生成的Verilog设计通过率提高12.7%,对于多模块任务提高13.7%,展示了VClare的规范修复能力以及大语言模型在前端硬件设计中的进一步潜力。
英文摘要
Large language models (LLMs) have demonstrated promising capabilities in generating Verilog code from natural language specifications. However, human-written specifications often contain semantic imperfections such as vagueness, contradictions, and incompleteness, which can significantly degrade the quality of hardware design generated by LLMs. In this paper, we present the first systematic study of imperfect specifications and propose an automated framework {VClare} to repair them to enhance the quality of resulting Verilog design. The proposed framework explores two complementary repair paradigms. The \textit{Spec-Level Repair} conducts LLM-driven inconsistency mining directly on the specification texts, while the \textit{Sim-Level Repair} employs simulation-based behavioral clustering with optional test-time inconsistency arbitration. In addition, we propose two new benchmark datasets with systematically injected specification defects. The first benchmark dataset is derived from the VerilogEval-human benchmark targeting single-module tasks, while the other benchmark dataset is derived from the ComplexVDB dataset and contains 53 multi-module tasks that reflect more realistic engineering scenarios. For single-module tasks, the {VClare} framework can repair the imperfections in the specifications effectively and thus enhance the pass rate of the generated Verilog design by 12.7\%, while for the multi-module tasks this enhancement can reach 13.7\%, demonstrating the capabilities of specification repair by {VClare} as well as further potential of LLMs in front-end hardware design.\footnote{The two benchmark datasets are released at https://anonymous.4open.science/r/VClare/.