arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM生成的SystemVerilog断言对保持语义的RTL变换的鲁棒性

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

FNU Aditi

arXiv 2609.05658首次发表:更新:

AI 中文总结

本文通过蜕变测试评估LLM生成SystemVerilog断言在保持语义的RTL变换下的鲁棒性,发现9.7%-27.0%的正确行为会失效,表明点准确性不足以衡量可靠性。

AI 中文摘要

大型语言模型(LLMs)正越来越多地被探索用于自动化SystemVerilog断言(SVA)生成,然而大多数评估仅报告在输入的单一语法表示上的正确性。这种点准确性并不能揭示当相同的RTL行为以不同方式编写时,模型的正确输出是否保持稳定。本文提出了一种在保持语义的RTL变换下对基于LLM的SVA生成进行受控的蜕变评估。从VERT数据集出发,我们构建了一个质量过滤的条件控制池和一个包含295个赋值行为的分层40程序评估集。我们评估了两个开源代码模型,Qwen2.5-Coder-7B和DeepSeek-Coder-V2-Lite,使用相同的评估提示和贪婪解码。研究了三种变换:操作数重排序、确定性标识符重命名和冗余括号化。除了基线和变换后的准确性外,我们还在RTL程序级别测量了条件鲁棒性、不变性失败和任意翻转率,并采用10,000样本的聚类自助法区间。在所有六个模型-变换条件下,9.7%-27.0%在原始RTL上正确的行为在保持语义的变换后变得不正确。因此,总体准确性可能掩盖了实质性的不稳定性:在标识符重命名下,DeepSeek-Coder-V2-Lite的准确性从53.9%提高到63.7%,而其原始正确行为中有19.5%失败。对30个采样到的正确到错误转换的人工审查识别出路径谓词丢失、分支极性错误、布尔结构损坏和输出契约违反。结果表明,仅靠点准确性不足以表征LLM在断言生成中的可靠性,并激励了面向鲁棒性的AI辅助硬件验证评估。

英文摘要

Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is written differently. This paper presents a controlled metamorphic evaluation of LLM-based SVA generation under semantics-preserving RTL transformations. Starting from the VERT dataset, we construct a quality-filtered conditional-control pool and a stratified 40-program evaluation set containing 295 assignment behaviors. We evaluate two open code models, Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, with an identical evaluation prompt and greedy decoding. Three transformations are studied: operand reordering, deterministic identifier renaming, and redundant parenthesization. Beyond baseline and transformed accuracy, we measure conditional robustness, invariance failure, and any-flip rate, with 10,000-sample clustered bootstrap intervals at the RTL-program level. Across all six model-transformation conditions, 9.7%-27.0% of behaviors that were correct on the original RTL become incorrect after a semantics-preserving transformation. Aggregate accuracy can therefore hide substantial instability: under identifier renaming, DeepSeek-Coder-V2-Lite improves from 53.9% to 63.7% accuracy while 19.5% of its originally correct behaviors fail. Manual review of 30 sampled correct-to-wrong transitions identifies dropped path predicates, branch-polarity errors, Boolean-structure corruption, and output-contract violations. The results show that point accuracy alone is insufficient for characterizing LLM reliability in assertion generation and motivate robustness-aware evaluation for AI-assisted hardware verification.

Comments10 pages, 3 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑