arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.04699cs.SEcs.AIcs.CL

R2Vul:通过强化学习与结构化推理蒸馏学习推理软件漏洞

R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation

Martin Weyssow, Chengran Yang, Junkai Chen, Ratnadira Widyasari, Ting Zhang, Huihui Huang, Huu Hung Nguyen, Yan Naing Tun, Tan Bui, Yikun Li, Ang Han Wei, Frank… 展开作者

Martin Weyssow, Chengran Yang, Junkai Chen, Ratnadira Widyasari, Ting Zhang, Huihui Huang, Huu Hung Nguyen, Yan Naing Tun, Tan Bui, Yikun Li, Ang Han Wei, Frank Liauw, Eng Lieh Ouh, Lwin Khin Shar, David Lo

更新

AI总结:

R2Vul通过RLAIF和结构化推理蒸馏训练小型代码LLM,在漏洞检测中生成有依据的解释,1.5B模型超越32B教师和商业LLM,并降低假阳性率。

AI中文摘要:

大型语言模型(LLMs)在软件漏洞检测方面展现出有前景的性能,但其推理能力仍不可靠。我们提出R2Vul,一种结合人工智能反馈强化学习(RLAIF)与结构化推理蒸馏的方法,用于训练小型代码LLM在生成安全感知解释的同时检测漏洞。与先前的思维链和指令微调方法不同,R2Vul通过RLAIF奖励有充分依据而非看似合理的漏洞解释,从而实现更精确的检测和高质量的推理生成。为支持RLAIF,我们构建了首个用于漏洞检测的多语言偏好数据集,包含C#、JavaScript、Java、Python和C语言的18,000个高质量样本。我们在五种编程语言上评估R2Vul,并与四种静态分析工具、八种最先进的基于LLM的基线以及各种微调方法进行对比。结果表明,1.5B参数的R2Vul模型超越了其32B教师模型以及Claude-4-Opus等领先商业LLM的性能。此外,我们引入了一个轻量级校准步骤,可在不同不平衡数据分布下降低假阳性率。最后,通过定性分析,我们证明LLM和人类评估者均一致认为R2Vul模型的推理优于其他基于推理的基线。

英文摘要:

Large language models (LLMs) have shown promising performance in software vulnerability detection, yet their reasoning capabilities remain unreliable. We propose R2Vul, a method that combines reinforcement learning from AI feedback (RLAIF) and structured reasoning distillation to teach small code LLMs to detect vulnerabilities while generating security-aware explanations. Unlike prior chain-of-thought and instruction tuning approaches, R2Vul rewards well-founded over deceptively plausible vulnerability explanations through RLAIF, which results in more precise detection and high-quality reasoning generation. To support RLAIF, we construct the first multilingual preference dataset for vulnerability detection, comprising 18,000 high-quality samples in C\#, JavaScript, Java, Python, and C. We evaluate R2Vul across five programming languages and against four static analysis tools, eight state-of-the-art LLM-based baselines, and various fine-tuning approaches. Our results demonstrate that a 1.5B R2Vul model exceeds the performance of its 32B teacher model and leading commercial LLMs such as Claude-4-Opus. Furthermore, we introduce a lightweight calibration step that reduces false positive rates under varying imbalanced data distributions. Finally, through qualitative analysis, we show that both LLM and human evaluators consistently rank R2Vul model's reasoning higher than other reasoning-based baselines.

↑