arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26751cs.LG

EquivSVA:跨等价RTL实现的行为断言的形式化验证数据集

EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations

FNU Aditi

首次发表
浏览论文内容

中文总结 AI 辅助

提出EquivSVA数据集,包含120个行为族、480个RTL实现和914个黄金属性,用于研究断言生成是否依赖实现细节,并通过Qwen2.5-Coder-7B模型评估展示其应用价值。

中文摘要 AI 辅助

大型语言模型越来越多地被用于从自然语言规范和寄存器传输级设计生成SystemVerilog断言。现有的数据集和基准支持重要目标,如大规模训练、形式化评估、规范到断言的生成以及基于突变的测试。一个补充性的需求是研究生成的断言是否捕获了外部可观察行为,还是依赖于某个RTL实现的偶然细节。我们提出了EquivSVA,一个围绕行为族组织的形式化验证数据集。每个族包含四个结构不同但具有相同外部可观察行为的RTL实现、共享的接口级黄金属性、三个受控突变体以及形式化验证证据。EquivSVA包含12个类别中的120个行为族、480个参考RTL实现、914个黄金属性和360个突变体。每个最终的族都通过一个固定的17任务验证套件,涵盖RTL等价性、黄金属性证明、属性可达性、突变体可区分性以及突变体上的黄金属性检查。我们还提供了固定的族安全训练、开发和测试划分。作为数据集所支持分析的一个小型演示,我们在保留的测试划分上评估了公开发布的、Apache-2.0许可的Qwen2.5-Coder-7B-Instruct模型。在293个仅接口生成的属性中,93个在形式上是健全的,并且在24个测试族中的14个中,健全属性的数量在不同等价实现之间有所变化。这些结果说明了行为族组织如何在不改变预期功能的情况下支持断言生成鲁棒性的受控研究。数据集、生成器、验证脚本和案例研究工件已在此https URL公开。

英文摘要

Large language models are increasingly used to generate SystemVerilog Assertions from natural-language specifica- tions and register-transfer-level designs. Existing datasets and benchmarks support important goals such as large- scale training, formal evaluation, specification-to-assertion generation, and mutation-based testing. A complemen- tary need is to study whether a generated assertion cap- tures externally observable behavior or depends on inci- dental details of one RTL implementation. We present EquivSVA, a formally verified dataset organized around behavior families. Each family contains four structurally distinct RTL implementations of the same externally ob- servable behavior, shared interface-level gold properties, three controlled mutants, and formal-validation evidence. EquivSVA contains 120 behavior families across 12 cat- egories, 480 reference RTL implementations, 914 gold properties, and 360 mutants. Every final family passes a fixed 17-job validation suite covering RTL equivalence, gold-property proofs, property reachability, mutant dis- tinguishability, and gold-property checks on mutants. We also provide fixed family-safe train, development, and test splits. As a small demonstration of the analyses en- abled by the dataset, we evaluate the publicly released, Apache-2.0-licensed Qwen2.5-Coder-7B-Instruct model on the held-out test split. Of 293 interface-only generated properties, 93 are formally sound, and the number of sound properties varies across equivalent implementations for 14 of 24 test families. These results illustrate how behavior-family organization can support controlled stud- ies of assertion-generation robustness without requiring changes in intended functionality. The dataset, generators, validation scripts, and case-study artifacts are publicly released at https://github.com/aditigupta96/EquivSVA.

补充信息

↑