arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VulnGym:面向代码智能体的仓库级漏洞检测基准测试

VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection

Kexing Ji, Jiachen Liu, Enze Hu, Cuiyun Gao, Keke Lian, Yongheng Liu, Lei Zhang, Tian Dong, Hao Chen, Wang Bin

arXiv 2608.02001首次发表:更新:

AI 中文总结

研究人员提出VulnGym基准,解决现有基准难以评估代码智能体仓库级漏洞检测能力的问题,发现当前代码智能体在该检测及支撑追踪构建上仍存在局限。

AI 中文摘要

基于大语言模型的漏洞检测近期取得了良好进展,而代码智能体进一步将该能力从孤立代码片段扩展至完整代码仓库。这一转变要求智能体自主探索仓库并定位与漏洞相关的代码,而非对预先选定的函数执行检测。然而,现有基准主要聚焦于对预先选定代码片段的漏洞分类,限制了其在仓库级漏洞检测中评估代码智能体的能力。此外,若缺乏细粒度的漏洞追踪标注,检测过程背后的能力局限仍难以探究。为解决这些局限,我们提出VulnGym,这是一个用于评估代码智能体漏洞检测能力的真实仓库级基准。VulnGym将已审核的GitHub安全公告与其对应的易受攻击版本仓库进行匹配,涵盖23个仓库中的184个安全公告和408个漏洞条目,每个条目都标注了行级入口点、关键操作和漏洞追踪。利用这些细粒度的真实标注,VulnGym定义了端到端检测任务和三个基于神谕的子任务,以联合评估漏洞检测并诊断代码定位和证据构建中的局限。我们的评估表明,当前代码智能体在端到端仓库级漏洞检测以及准确支撑追踪的构建方面仍存在局限。

英文摘要

Recent advances in LLM-based vulnerability detection have shown promising results, while coding agents further extend this capability from isolated code snippets to complete repositories. This shift requires agents to autonomously explore repositories and locate vulnerability-relevant code, instead of performing detection on preselected functions. However, existing benchmarks primarily focus on vulnerability classification over preselected code snippets, limiting their ability to evaluate coding agents in repository-level vulnerability detection. Moreover, without fine-grained vulnerability trace annotations, the capability limitations underlying the detection process remain difficult to explore. To address these limitations, we present \textbf{VulnGym}, a real-world repository-level benchmark for evaluating vulnerability detection by coding agents. VulnGym aligns reviewed GitHub advisories with their corresponding vulnerable version repositories. It contains 184 advisories and 408 vulnerability entries across 23 repositories, with each entry annotated with line-level entry points, critical operations, and vulnerability traces. Using this fine-grained ground truth, VulnGym defines an end-to-end detection task and three oracle-based subtasks to jointly evaluate vulnerability detection and diagnose limitations in code localization and evidence construction. Our evaluation indicates that current coding agents remain limited in both end-to-end repository-level vulnerability detection and the construction of accurate supporting traces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑