arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VeriBugBench:一个基于经验构建Verilog RTL调试基准的框架

VeriBugBench: An Empirically Grounded Framework for Constructing Verilog RTL Debugging Benchmarks

Xiankai Meng, Kejian Feng, Xinlin Zhao, Zhuo Zhang, Yan Lei, Xiaoguang Mao, Jiang Wu

arXiv 2609.18022首次发表:更新:

发表机构

School of Computer and Information Engineering, Institute for Artificial Intelligence, Shanghai Polytechnic University; Guangzhou College of Commerce; Chongqing University; National University of Defense Technology; Academy of Military Sciences(上海理工大学计算机与信息工程学院人工智能研究院; 广州商学院; 重庆大学; 国防科技大学; 军事科学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VeriBugBench提出一个基于经验构建Verilog RTL调试基准的框架,通过故障构建、LLM测试增强和保留机制,生成2,608个可执行故障实例,提升故障可观测性至39.54%。

AI 中文摘要

RTL源代码级调试研究需要基准工件,这些工件需提供带有精确变更位置的有缺陷设计、可执行的测试激励以及可复现的配置。现有的Verilog资源通常仅提供这些元素中的一部分。我们提出了VeriBugBench,一个通过基于经验的故障构建、基于LLM的测试平台增强和基于执行的保留来构建Verilog RTL调试基准的框架。变异库将RTL缺陷修复历史中观察到的重复性、多粒度修复模式映射到19个可执行的逆算子。对于每个项目,LLM从干净的DUT和原始测试平台生成一个设计特定的激励阶段;该阶段与原始测试平台组合以进行候选执行。将该框架应用于45个开源项目,生成了VeriBugBench-v1.0,包含2,608个可执行的单故障实例,其影响可在设计输出端观察到。在这45个项目中,组装后的测试平台将平均项目级故障可观测性从36.01%提高到39.54%,并平均改善了行覆盖率和执行轨迹多样性。VeriBugBench提供了版本化的RTL变体、源代码级真实标签、测试平台和执行工件,用于评估RTL调试方法。

英文摘要

RTL source-level debugging research requires benchmark artifacts that provide faulty designs together with precise change locations, executable test stimuli, and reproducible configurations. Available Verilog resources usually provide only a subset of these elements. We present VeriBugBench, a framework for constructing Verilog RTL debugging benchmarks through empirically grounded fault construction, LLM-based testbench enhancement, and execution-based retention. The mutation library maps recurring, multi-granularity repair patterns observed in RTL bug-fix histories to 19 executable inverse operators. For each project, an LLM generates a design-specific stimulus phase from the clean DUT and original testbench; the phase is composed with the original testbench for candidate execution. Applying the framework to 45 open-source projects yields VeriBugBench-v1.0, with 2,608 executable single-fault instances whose effects are observable at design outputs. Across the 45 projects, the assembled testbenches increase mean project-level fault observability from 36.01% to 39.54% and improve line coverage and execution-trace diversity on average. VeriBugBench provides versioned RTL variants, source-level ground truth, testbenches, and execution artifacts for evaluating RTL debugging methods.

Comments14 pages, 4 figures, and 6 tables. Submitted to IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD). This manuscript substantially extends the DAC 2023 paper "MANTRA: Mutation Testing of Hardware Design Code Based on Real Bugs" (DOI: 10.1109/DAC56929.2023.10247962)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑