发表机构
Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示LLM漏洞修复基准在智能体、框架和数据集三个层面存在可靠性陷阱,通过112个真实漏洞实验证明基准分数高度依赖评估设计,且模型修复质量与开发者意图仍有差距。
AI 中文摘要
大型语言模型(LLMs)在自动化漏洞修复方面展现出巨大潜力,但当前的基准测试可能会显著扭曲所报告的性能。基于我们在开发、运行和压力测试此类框架方面的丰富经验,我们从三个维度识别了未被充分研究的陷阱:(1)智能体层面的因素,其中提示设计、工具可用性和详细指令可以在不提升开发者对齐的补丁质量的情况下提高成功率;(2)框架层面的因素,其中权限错误、基础设施缺陷和超时处理可能悄无声息地抑制或夸大性能;(3)数据集层面的因素,其中错误报告和单一概念验证(PoC)测试无法捕捉补丁是否解决了根本原因或遵循开发者意图。我们从84个开源C/C++、Go和Rust项目中精选了112个历史漏洞,每个漏洞都配有PoC测试、回归测试以及评估与原始开发者设计原则一致性的额外开发者测试。通过受控实验和案例研究,我们表明LLMs在理想条件下可以达到较高的PoC通过率,但基准执行选择可能实质性改变测量的成功率。更重要的是,开发者测试通过率仍然较低,且随着新模型的推出仅略有提升,这表明模型越来越多地抑制症状,而没有持续产生上游质量的修复。这些结果表明,基准分数对评估设计高度敏感,我们为更严格、可靠和可复现的评估提供了实用指南。
英文摘要
Large language models (LLMs) have shown strong potential for automated vulnerability patching, but current benchmarks can substantially distort reported performance. Drawing on extensive experience developing, running, and stress-testing such frameworks, we identify under-examined pitfalls across three dimensions: (1) agent-level factors, where prompting, tool availability, and detailed instructions can raise success rates without improving developer-aligned patch quality; (2) framework-level factors, where permission errors, infrastructure bugs, and timeout handling can silently suppress or inflate performance; and (3) dataset-level factors, where bug reports and single proof-of-concept (PoC) tests fail to capture whether patches address root causes or follow developer intent. We curate 112 historical bugs from 84 open-source C/C++, Go, and Rust projects, each with PoC tests, regression tests, and additional developer tests that assess alignment with the original developers' design principles. Through controlled experiments and case studies, we show that LLMs can achieve high PoC passing rates under ideal conditions, yet benchmark execution choices can materially change measured success. More importantly, developer-test passing rates remain low and improve only marginally with newer models, suggesting that models increasingly suppress symptoms without consistently producing upstream-quality fixes. These results show that benchmark scores are highly sensitive to evaluation design, and we provide practical guidelines for more rigorous, reliable, and reproducible evaluation.