VSpector:面向RISC-V CPU的规范驱动缺陷检测
VSpector: Specification-Driven Bug Detection for RISC-V CPUs
- Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
VSpector利用官方RISC-V规范,通过四阶段流水线自动检测CPU RTL实现违规,在CVA6和XiangShan上发现42个新缺陷,优于现有模糊测试工具。
AI中文摘要:
检测开源RISC-V CPU实现中的RTL设计缺陷对于确保系统可靠性至关重要。传统检测方法本质上依赖于预定义的工件。在本文中,我们利用官方的、自然语言的RISC-V规范作为缺陷检测的有效信息来源。我们提出了VSpector,一个规范驱动的缺陷检测流水线,它直接检查CPU寄存器传输级(RTL)实现是否符合官方规范规则,而无需专门构建参考模型、形式化属性或自定义缺陷模式。为了解决使用大型语言模型(LLM)时广泛上下文范围与模型推理准确性之间的关键技术权衡,VSpector在四阶段流水线中采用了逐步上下文细化方案:规则提取、实现定位、候选识别和顺序违规审计。我们在两个工业级RISC-V CPU(CVA6和XiangShan)上评估了VSpector。在报告的217个候选中,人工检查确认了148个真实违规,精确率为68.2%。这些违规对应73个不同的缺陷,包括42个先前未知的缺陷。在我们的对比实验中,最先进的CPU模糊测试工具DiveFuzz在每CPU运行24小时内未检测到这些新缺陷中的任何一个。所有42个新缺陷均已向上游报告,开发者已修复19个并确认另外11个(共30个),这表明规范驱动的审计是CPU缺陷检测的一种实用且互补的策略。
英文摘要:
Detecting RTL design bugs in open-source RISC-V CPU implementations is critical for ensuring system reliability. Traditional detection approaches inherently rely on predefined artifacts. In this paper, we leverage the official,natural-language RISC-V specifications as an effective information source for bug detection. We present VSpector, a specification-driven bug detection pipeline that directly checks whether CPU register-transfer level (RTL) implementations adhere to official specification rules, without requiring specialized construction of reference models, formal properties, or custom bug patterns. To resolve the key technical trade-off between broad context scope and model reasoning accuracy when using Large Language Models (LLMs), VSpector employs a stepwise context refinement scheme across a four-stage pipeline: rule extraction, implementation localization, candidate identification, and sequential violation auditing. We evaluate VSpector on two industrial-strength RISC-V CPUs, CVA6 and XiangShan. Out of 217 reported candidates, manual inspection confirmed 148 true violations, representing a 68.2% precision. These violations correspond to 73 distinct bugs, including 42 previously unknown bugs. In our comparative experiments, DiveFuzz, a state-of-the-art CPU fuzzer, detected none of these new bugs during 24-hour runs per CPU. All 42 new bugs have been reported upstream, with developers already fixing 19 and confirming an additional 11 (30 in total), demonstrating that specification-driven auditing is a practical and complementary strategy for CPU bug detection.