发表机构
Singapore Management University; Harbin Institute of Technology; GovTech(新加坡管理大学; 哈尔滨工业大学; 新加坡政府科技局)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有漏洞检测孤立处理函数的问题,提出基于CPG的智能体强化学习框架VulAgentRL,通过图验证设计可靠奖励并预热初始化策略,在仓库级划分下的严格指标及多场景中均优于基线。
AI 中文摘要
现实世界中的漏洞往往跨越多个函数,但大多数基于学习的检测器会孤立地对每个函数进行分类:在对真实CVE样本的分析中,我们发现71.7%的漏洞函数需要函数外部的证据才能被正确分类。智能体强化学习(RL)可通过让模型自行收集此类证据来缩小这一差距,但它缺乏可靠的奖励机制,因为仅基于最终判决定义的奖励可能在未开展任何调查的情况下获得。我们提出VulAgentRL,这是一个基于代码属性图(CPG)的过程间漏洞检测智能体强化学习框架。CPG兼具两项作用:推理时,策略会向其查询调用者、被调用者、数据流及其他信息;训练时,同一图会验证策略所引用的证据。由于每个CPG节点都带有持久整数标识符,该验证为精确比对而非文本匹配,因此奖励会将判决归因于有证据支持的情况。我们还通过蒸馏教师调查来初始化策略,并证明这种预热启动是必要的,因为RL无法获取其从未采样过的工具使用行为。在防止数据泄露的仓库级划分下,VulAgentRL在严格的两两正确指标上优于包括前沿模型在内的现有最佳基线,同时发出更少的工具调用,且其优势在分布外语料库和类别不平衡场景下依然存在。
英文摘要
Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sample of real CVEs, we find that 71.7% of vulnerable functions require evidence from outside the function to be classified correctly. Agentic reinforcement learning (RL) could close this gap by enabling a model to gather that evidence itself, but it lacks a reliable reward, since a reward defined on the final verdict alone can be obtained without performing any investigation. We propose VulAgentRL, an agentic RL framework for interprocedural vulnerability detection built on a Code Property Graph (CPG). The CPG serves two roles: at inference time the policy queries it for callers, callees, dataflow, and other queries, and at training time the same graph verifies the evidence the policy cites. Because every CPG node carries a persistent integer identifier, this verification is an exact comparison rather than a textual match, so the reward credits verdicts that are supported by evidence. We further initialize the policy by distilling teacher investigations, and show that this warm start is necessary, since RL cannot acquire tool-use behavior it never samples. Under a repository-level split that prevents leakage, VulAgentRL outperforms state-of-the-art baselines, including frontier models, on the strict pair-wise-correct metric while issuing fewer tool calls, and its advantage persists on an out-of-distribution corpus and under class imbalance.