arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对用于Verilog设计流程的大语言模型进行基准测试

Benchmarking LLMs for Verilog Design Flows

Angshuman Chakravertty, Rahul Koshti, Buddhi Prakash Sharma, Vinay Chamola

arXiv 2607.22759首次发表:更新:

发表机构

School of Technology Management and Engineering, NMIMS; Indian Institute of Technology, Roorkee; Birla Institute of Technology and Science, Pilani (BITS Pilani)(NMIMS技术管理与工程学院; 鲁尔基印度理工学院; 皮拉尼贝拉理工学院(BITS Pilani))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型生成Verilog RTL代码能力的基准测试问题,提出含特定流程的可重复平台,经测试提升了开源模型的语法有效性和模拟通过率,且平台与数据集开源,利于生成式人工智能在硬件设计工作流的评估。

AI 中文摘要

大语言模型在代码生成方面展现出潜力,但生成正确、可合成的硬件描述语言(HDL)代码的能力仍有待适当基准测试。现有评估主要依赖pass@k指标且缺乏端到端工具链验证。本文提出一个可重复的基准测试平台,在50个精心策划的任务上评估开源大语言模型的Verilog RTL生成。通过特定流程验证生成代码,在评估三个不同大小模型的12个基准测试和1610次总运行中,提高了语法有效性和模拟通过率。平台和数据集开源,可对硬件设计工作流的生成式人工智能进行可重复评估。

英文摘要

Large language models (LLMs) show promise in code generation, but their capabilities to produce correct, synthesizable hardware description language (HDL) code still remain to be properly benchmarked. Existing evaluations are primarily relying on pass@k metrics and lack proper end-to-end toolchain validation. This paper presents a reproducible benchmarking platform that evaluates open-source LLMs on Verilog RTL generation across 50 curated tasks consisting of combinational, sequential, finite state machine (FSM), and mixed designs. The pipeline consisting of constrained prompting, post-processing, and semantic-aware iterative refinement with waveform analysis, formal equivalence verification, and Abstract Syntax Tree (AST)-based repair validates the generated code via Verilator compilation and Icarus Verilog simulation. Across the 12 benchmarks and the 1,610 total runs evaluating three models of different sizes (Llama-3-8B, StarCoder2-7B, and TinyLlama-1.1B), the pipeline improved syntax validity from 0% to a 70.43% average and simulation pass rate to 51.8% across three open-source models. Most notably TinyLlama (1.1B parameters) achieved the highest individual syntax validity at 80.0%, with functional correctness comparable to the 8B model. The platform and dataset are open-source, enabling reproducible evaluation of generative AI for hardware design workflows.

Comments7 pages, 3 figures, 3 tables. Manuscript prepared for submission to IEEE Design & Test

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑