arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05380cs.PLcs.ARcs.CL

VHDL-REPOBENCH:用于评估大型语言模型在VHDL设计生成能力的仓库级基准

VHDL-REPOBENCH: A Repository-Level Benchmark for Evaluating Large Language Models on VHDL Design Generation

  • IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

Prashanth Vijayaraghavan, Akul Malhotra, Ashutosh Jadhav, Ehsan Degan, Vandana Mukherjee

AI总结:

针对VHDL设计生成缺乏仓库级基准的问题,提出VHDL-REPOBENCH,包含约100个仓库、2.5k文件与500测试平台,评估多模型,发现多文件推理与层次设计仍是挑战。

AI中文摘要:

大型语言模型(LLM)越来越多地应用于硬件设计自动化,展现出在生成和理解硬件描述语言方面的强大潜力。然而,现有的大多数基准侧重于Verilog,对VHDL的评估有限,而VHDL在工业界和学术界仍广泛用于FPGA和安全关键系统。为弥补这一空白,我们引入了VHDL-REPOBENCH,一个大规模、跨文件、仓库级的基准,用于评估LLM在真实VHDL设计生成和分析任务中的能力。VHDL-REPOBENCH精选了约100个开源VHDL仓库,涵盖约2.5k个VHDL文件和约500个测试平台,并提供结构化的问题陈述、模块存根和自验证测试平台。该基准支持对语法、语义正确性、层次推理、跨文件依赖解析和功能验证进行全面评估。我们评估了多个最先进的模型,包括GPT-4o、Llama-3-70B、Qwen2.5-72B、CodeLlama-70B,以及多步推理方法如Reflexion和CoDes。结果表明,尽管当前LLM在行级和块级准确性上表现中等,但在多文件推理、层次化设计理解和规范到模块生成方面仍面临重大挑战。VHDL-REPOBENCH是首个大规模VHDL专用基准,为硬件设计社区评估、比较和推进LLM在实际VHDL开发中的能力提供了宝贵资源。

英文摘要:

Large Language Models (LLMs) are increasingly applied in hardware design automation, demonstrating strong potential in generating and understanding hardware description languages. However, most existing benchmarks focus on Verilog, with limited evaluation of VHDL, which remains widely used in industry and academia for FPGA and safety-critical systems. To address this gap, we introduce VHDL-REPOBENCH, a large-scale, cross-file, repository-level benchmark for assessing LLM capabilities on realistic VHDL design generation and analysis tasks. VHDL-REPOBENCH curates ~100 open-source VHDL repositories, encompassing ~2.5k VHDL files and ~500 testbenches, and provides structured problem statements, module stubs, and self-verifying testbenches. The benchmark enables comprehensive evaluation across syntax, semantic correctness, hierarchical reasoning, cross-file dependency resolution, and functional verification. We evaluate several state-of-the-art models, including GPT-4o, Llama-3-70B, Qwen2.5-72B, CodeLlama-70B, and multi-step reasoning approaches such as Reflexion and CoDes. Results reveal that while current LLMs achieve moderate line- and block-level accuracy, substantial challenges remain in multi-file reasoning, hierarchical design understanding, and specification-to-module generation. VHDL-REPOBENCH represents the first large-scale VHDL-focused benchmark and provides a valuable resource for the hardware design community to evaluate, compare, and advance LLM capabilities for practical VHDL development.

补充信息

↑