发表机构
Esslingen University; University of Tübingen(埃斯林根应用技术大学; 图宾根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过复现六个开源RAG4SVD系统、建立统一基准并进行组件级分析,揭示检索增强漏洞检测的可复现性与性能依赖问题,强调需在流水线组件层面进行评估。
AI 中文摘要
检索增强生成(RAG)越来越多地被用于增强基于大型语言模型(LLM)的软件漏洞检测,通过将预测基于检索到的漏洞知识(如漏洞报告)来实现。然而,现有的基于RAG的软件漏洞检测(RAG4SVD)系统通常使用专有模型进行评估,这对开放科学和可复现性提出了挑战。此外,研究使用了不同的数据集、自定义知识库、不同的骨干模型和多样的指标,这阻碍了有意义的跨系统比较。在本工作中,我们研究了六个开源RAG4SVD系统,并通过以下方式解决这些可复现性和可比性挑战:(i)在开放权重设置下复现其实验设置,以及(ii)使用公共数据集、指标套件和开放权重模型池建立统一基准。此外,RAG4SVD系统通常由多个组件组成,但往往仅作为整体系统(即端到端)进行评估。因此,我们进行了(iii)组件级分析,将具有代表性的RAG4SVD流水线分解为输入抽象、知识检索和检测。我们的结果表明,各系统的可复现性差异很大。在所提出的统一基准下,已发表的RAG4SVD性能在受控的开放权重评估下无法迁移,并且强烈依赖于所使用的模型。组件分析表明,有效的RAG4SVD依赖于流水线各阶段之间的对齐。例如,oracle知识将检索提升至接近最优,但性能仍然较低(成对准确率为0.51),这表明仅靠检索有效性不足以实现可靠检测。这些发现促使不仅在端到端层面,而且在流水线组件层面评估RAG4SVD,并为更标准化、RAG感知的评估实践提供了基础。
英文摘要
Retrieval-Augmented Generation (RAG) is increasingly used to enhance Large Language Model (LLM)-based software vulnerability detection by grounding predictions in retrieved vulnerability knowledge, such as vulnerability reports. However, existing RAG-based software vulnerability detection (RAG4SVD) systems are often evaluated using proprietary models, which challenges open science and reproducibility. Further, studies use different datasets, custom knowledge bases, different backbone models, and diverse metrics, which hinders meaningful cross-system comparison. In this work, we study six open-source RAG4SVD systems and address these reproducibility and comparability challenges through (i) reproduction of their experimental settings under an open-weight setting, and (ii) a unified benchmark using a common dataset, metric suite, and pool of open-weight models. Further, RAG4SVD systems typically consist of multiple components, yet are often evaluated only as a whole system, i.e., end-to-end. Therefore, we perform (iii) a component-level analysis that decomposes representative RAG4SVD pipelines into input abstraction, knowledge retrieval, and detection. Our results demonstrate that reproducibility varies substantially across systems. Under the presented unified benchmark, published RAG4SVD performance does not transfer under a controlled open-weight evaluation and depends strongly on the used model. The component analysis shows that effective RAG4SVD depends on the alignment between pipeline stages. For example, oracle knowledge raises retrieval to near-optimal, yet performance remains low (0.51 pairwise accuracy), demonstrating that retrieval effectiveness alone is insufficient for reliable detection. These findings motivate evaluating RAG4SVD not only end-to-end, but at the level of pipeline components, and provide a basis for more standardized, RAG-aware evaluation practices.