AI 中文总结
SWE-Test通过输入预测任务,利用覆盖率引导模糊测试挖掘真实C/C++程序中的深层分支,将漏洞发现分解为三种模式,评估发现约束推断是主要瓶颈,最佳模型通过率仅55%。
AI 中文摘要
漏洞发现正成为大语言模型(LLM)智能体的一项重要能力:那些默默遗漏真实缺陷的智能体会使关键软件暴露于风险之中。因此,严格衡量这一能力迫在眉睫,但现有基准存在数据污染风险,对未知漏洞集的召回率评分可被操纵,且往往依赖合成漏洞,并仅给出单一端到端结果,无法定位智能体失败的具体环节。漏洞发现是一项复合能力:智能体必须理解源代码、推断输入约束、构造输入、执行输入,并根据反馈迭代修正。我们将其衡量重新构建为一项具有封闭、确定性真值的输入预测任务:利用覆盖率引导的模糊测试,我们在真实世界的C/C++程序中挖掘深层目标分支,并要求智能体预测一个能将执行驱动到给定分支的输入。这将漏洞发现分解为三种任务模式,覆盖跨越15个领域的22个真实世界C/C++程序。开环模式和带反馈模式共享16个代码库(13个领域)中的60个固定目标任务实例,分别测试无距离预言机与有距离预言机情况下的输入构造能力,以将代码理解与反馈驱动的修正相分离。在线竞技场模式则移除预定义目标,根据覆盖率增益对路径探索进行评分,使用一个包含11个程序的独立且部分重叠的池;智能体共同在六个程序中确认了13个不同的缺陷。评估15种默认努力程度的模型-脚手架配置,最佳配置在带反馈模式下仅达到55.0%的通过率,而七组配对的Claude Code配置在带反馈情况下的平均通过率为36.4%,无反馈时为19.3%。对失败进行分解,我们发现约束推断(而非导航)是主要的瓶颈。我们发布了SWE-Test,并附带一个即用型评估环境。
英文摘要
Vulnerability discovery is becoming an important ability of large language model (LLM) agents: agents that silently miss real defects leave critical software exposed. Rigorously measuring this ability is therefore urgent, but existing benchmarks are gameable through data contamination, score recall against an unknowable vulnerability set, often rely on synthetic bugs, and report a single end-to-end verdict that cannot localize where an agent fails. Vulnerability discovery is a composite ability: an agent must comprehend source code, infer input constraints, construct inputs, execute them, and iteratively correct from feedback. We recast its measurement as an input-prediction task with a closed, deterministic ground truth: using coverage-guided fuzzing, we mine deep target branches in real-world C/C++ programs and ask an agent to predict an input that drives execution to a given branch. This decomposes discovery into three task modes over 22 real-world C/C++ programs spanning 15 domains. Open-loop and Feedback-enabled share 60 fixed-target task instances across 16 of these codebases (13 domains), testing input construction without and with a distance oracle to isolate code comprehension from feedback-driven correction. Online Arena instead removes the predefined target and scores path exploration by coverage gain on a separate, partially overlapping pool of 11 programs; agents collectively confirmed 13 distinct bugs across six programs. Evaluating 15 default-effort model-scaffold configurations, the best reaches only 55.0% pass rate in the Feedback-enabled mode, and the mean across seven paired Claude Code configurations is 36.4% with feedback versus 19.3% without. Decomposing failures, we find constraint inference, not navigation, is the dominant bottleneck. We release SWE-Test with a turnkey evaluation environment.