arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16729cs.AR

SpecLens:基于行为分歧的规范派生约束的LLM Verilog生成

SpecLens: LLM-Based Verilog Generation with Specification-Derived Constraints via Behavioral Divergence

Wen Bing, Bing Li

首次发表
浏览论文内容

中文总结 AI 辅助

SpecLens通过分析多个候选实现的行为分歧,从规范派生约束,提升LLM生成Verilog的功能正确性,在VerilogEval v2.0上达到86.2%-89.4%的pass@1。

中文摘要 AI 辅助

大型语言模型(LLM)近期在Verilog生成方面展现出潜力,但直接从自然语言规范生成功能正确的RTL仍然是一项极具挑战性的任务。现有方法主要通过检索增强生成(RAG)、自规划或少量样本提示来改进基于LLM的Verilog生成。然而,这些方法主要侧重于外部或通用形式的增强,而非通过任务特定约束强化规范。在本工作中,我们提出SpecLens,一个基于LLM的Verilog生成自动化框架,通过分析多个候选实现之间的行为分歧来派生规范驱动的约束,在生成过程中仅使用原始规范作为唯一外部语义来源。在VerilogEval v2.0规范到RTL基准上,SpecLens在o3-mini-medium上实现86.2%的功能pass@1比率,在o3-mini-high上实现89.4%。这对应于在o3-mini-medium上比SOTA提示方法高出3.6个百分点,在o3-mini-high上比SOTA行为分歧方法高出3.8个百分点。此外,在RTLLM v1.1和v2.0上,分析表明SpecLens更忠实于规范,且不易受基准对齐先验影响。SpecLens在VerilogEval v2.0上实现100%的语法正确性,在RTLLM v1.1上为86.2%,在RTLLM v2.0上为88%,即使不使用昂贵的编译修复循环来迭代修改生成的代码。代码开源,可在该https URL获取。

英文摘要

Large language models (LLMs) have recently shown promise in Verilog generation, but producing functionally correct RTL directly from natural-language specifications remains a highly challenging task. Existing approaches improve LLM-based Verilog generation mainly with retrieval-augmented generation (RAG), self-planning, or few-shot prompting. However, these methods focus primarily on external or generic forms of enhancement rather than strengthening the specification with task-specific constraints. In this work, we propose SpecLens, an automated framework for LLM-based Verilog generation that derives specification-driven constraints by analyzing behavioral divergence among multiple candidate implementations, using the original specification as the only external semantic source during generation. On the VerilogEval v2.0 spec-to-RTL benchmark, SpecLens achieves a functional pass@1 ratio of 86.2\% with o3-mini-medium and 89.4\% with o3-mini-high. This corresponds to a 3.6 percentage-point gain over the SOTA prompting method with o3-mini-medium and a 3.8 percentage-point gain over the SOTA behavioral divergence method with o3-mini-high. In addition, on RTLLM v1.1 and v2.0, analysis shows that SpecLens is more specification-faithful and less prone to benchmark-aligned priors. SpecLens achieves 100\% syntactic correctness on VerilogEval v2.0, 86.2\% on RTLLM v1.1, and 88\% on RTLLM v2.0, even without using costly compile-repair loops to revise generated code iteratively. The code is open source and available at https://anonymous.4open.science/r/SpecLens-4632/readme.md.

发表机构

  • TU Ilmenau(埃尔福特工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑