arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当模糊测试与理解相遇:用于RTL验证的大语言模型驱动的语义测试生成

When Fuzzing Meets Understanding: LLM-Driven Semantic Test Generation for RTL Verification

Kun Wang, Cangyuan Li, Kaiyan Chang, Siyang Cai, Yinhe Han, Ying Wang

arXiv 2607.10340首次发表:更新:

AI 中文总结

针对硬件验证难题,提出ChipFuzzer框架,利用大语言模型语义推理能力。采用双阶段工作流程,覆盖引导阶段结合控制流分析提升覆盖率,错误引导阶段借助历史数据提高错误发现率,实验显示其显著提升验证效果。

AI 中文摘要

现代芯片日益增长的复杂性给硬件验证带来了重大挑战。近年来,覆盖引导的模糊测试已成为提高验证效率的一种有前途的方法。然而,现有的硬件模糊器仍难以实现高覆盖率并发现极端情况的错误,因为它们主要依赖启发式策略,对被测设计(DUT)的内部逻辑和语义行为的推理能力有限。在这项工作中,我们提出了ChipFuzzer,这是一个硬件模糊测试框架,它利用大语言模型(LLMs)的语义推理能力来提高模糊测试的有效性。ChipFuzzer采用双阶段工作流程,包括覆盖引导阶段和错误引导阶段。在覆盖引导阶段,ChipFuzzer采用控制流相似性和差异分析来指导由LLM驱动的测试用例生成,从而提高覆盖率。在错误引导阶段,ChipFuzzer利用历史错误数据来识别容易出现错误的代码区域,并为这些区域的测试用例生成确定优先级,从而提高错误发现效率。在三个开源CPU设计上的实验结果表明,与最强的基线相比,ChipFuzzer将平均条件覆盖率提高了5.8个百分点,将错误检测率提高了21.1个百分点。

英文摘要

The growing complexity of modern chips poses significant challenges to hardware verification. In recent years, coverage-guided fuzzing has emerged as a promising approach for improving verification efficiency. However, existing hardware fuzzers still struggle to achieve high coverage and expose corner-case bugs, as they predominantly rely on heuristic strategies with limited ability to reason about the internal logic and semantic behavior of the design under test (DUT). In this work, we propose ChipFuzzer, a hardware fuzzing framework that leverages the semantic reasoning capabilities of large language models (LLMs) to improve fuzzing effectiveness. ChipFuzzer adopts a dual-stage workflow comprising a Coverage-Guided stage and a Bug-Guided stage. In the Coverage-Guided stage, ChipFuzzer employs control-flow similarity and discrepancy analysis to guide LLM-driven testcase generation, thereby improving coverage. In the Bug-Guided stage, ChipFuzzer leverages historical bug data to identify bug-prone code regions and prioritize testcase generation for those regions, thus enhancing bug discovery efficiency. Experimental results on three open-source CPU designs show that ChipFuzzer improves average condition coverage by 5.8 percentage points and bug detection rate by 21.1 percentage points over the strongest baseline.

Comments8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑