发表机构
Cisco Systems Inc.; Carnegie Mellon University; Yale University(思科系统公司; 卡内基梅隆大学; 耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出VLoc Bench基准,包含500个真实漏洞,评估27个语言模型和4个静态分析工具在仓库规模上的漏洞定位能力,发现最强系统F1仅0.229,且定位能力与修复后可靠性不相关。
AI 中文摘要
语言模型智能体越来越多地操作于完整的软件仓库,但网络安全评估主要衡量它们能否检测、复现或修复漏洞,而非能否定位相关代码。我们研究漏洞定位:给定一个弱点类别和一个不熟悉的仓库,识别与该弱点相关的实现文件。我们引入了漏洞定位基准(VLoc Bench),包含来自六个包生态系统和147个CWE类别的290个仓库中的500个真实世界漏洞。每个任务配对安全修复前后的仓库快照。在易受攻击的快照上,智能体仅接收CWE描述和只读终端访问权限,必须返回受影响的文件;在已修补的快照上,它必须确定记录的漏洞不再存在。我们在一个通用智能体接口下评估了27个语言模型和四个静态分析工具。仓库规模的漏洞定位仍然困难:最强系统达到0.229的文件F1分数,38.4%的任务未被任何评估模型正确定位。我们进一步发现,更强的定位能力并不意味修复后的可靠行为:能有效识别易受攻击文件的系统在已修补的仓库上仍可能报告不支持的定位。这些结果确立了漏洞定位作为一种独特的仓库规模能力,并为研究安全智能体如何搜索易受攻击代码以及何时应弃权(不执行)报告提供了环境。
英文摘要
Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant code. We study vulnerability localization: given a weakness class and an unfamiliar repository, identify the implementation files associated with that weakness. We introduce the Vulnerability Localization Benchmark (VLoc Bench), comprising 500 real world vulnerabilities from 290 repositories across six package ecosystems and 147 CWE categories. Each task pairs repository snapshots immediately before and after a security fix. On the vulnerable snapshot, an agent receives only the CWE description and read-only terminal access and must return the affected files; on the patched snapshot, it must determine that the recorded vulnerability is no longer present. We evaluate 27 language models and four static-analysis tools under a common agent interface. Repository-scale vulnerability localization remains difficult: the strongest system achieves 0.229 File F1, and 38.4% of tasks receive no correct localization from any evaluated model. We further find that stronger localization does not imply reliable behavior after remediation: systems that identify vulnerable files effectively can still report unsupported locations on patched repositories. These results establish vulnerability localization as a distinct repository-scale capability and provide a setting for studying both how security agents search for vulnerable code and when they should refrain from reporting it.
Comments29 pages, 6 figures, technical report for VLoc-bench