基于LLM的代码漏洞修复中的指标失效:一项实证研究与变更感知筛选
Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen
浏览论文内容
中文总结 AI 辅助
本研究通过五项受控实验证明编译率作为LLM漏洞修复指标不可靠,并提出变更感知的diff_F1作为廉价筛选工具,主张采用变更感知且基于执行的评估方法。
中文摘要 AI 辅助
大语言模型(LLMs)越来越多地被应用于C/C++安全漏洞的自动修复,而编译率是常被报告的进展代理指标:即生成的补丁能否编译通过。我们认为,编译率对于单函数漏洞修复而言是一个科学上不可靠的指标,我们通过五项受控实验来支持这一观点,实验涉及来自Big-Vul的203个易受攻击函数、三个开源代码LLM(参数量从3.5亿到67亿)以及三种提示策略。编译率(i)对一项能显著改善生成代码的干预措施几乎无响应;(ii)主要受评估框架和数据集伪影影响,而非模型质量,约64%的编译失败不能归因于模型,且这一比例在不同模型间几乎不变;(iii)在单个编译器标准标志下,相同补丁的编译率变化1.8至2.7倍,且无回归;(iv)对三个模型的排序与参考相似度指标相反;(v)当用作优化目标时,会奖励未修复的情况,因为编译器反馈循环提高了编译率,而与人工修复的相似度下降,人工检查发现新编译通过的输出中存在删除式和占位符式的未修复。自然的替代方案——全函数CodeBLEU——同样失败:未更改的易受攻击输入副本得分高于所有模型。我们还考察了diff_F1,一种仅对编辑区域评分的变更感知筛选。它对无操作给予零分,对我们观察到的部分(尽管不是全部)基于删除的作弊补丁给予接近零分,同时仍对真正的部分编辑给予肯定,因此它可能作为更深层、基于执行的分析之前的廉价筛选。它不是一个修复质量指标,我们报告了其不足之处。我们的研究结果主张对基于LLM的漏洞修复进行变更感知、基于执行的评估。
英文摘要
Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, and we support this with five controlled experiments over 203 vulnerable functions from Big-Vul, three open-source code LLMs (350M to 6.7B parameters), and three prompting strategies. Compile rate (i) barely responds to an intervention that substantially improves the generated code; (ii) is dominated by evaluation-harness and dataset artifacts rather than model quality, with about 64% of compile failures not attributable to the model, a share that is nearly invariant across models; (iii) shifts by 1.8 to 2.7 times on identical patches under a single compiler-standard flag, with zero regressions; (iv) ranks the three models in the opposite order to reference-similarity metrics; and (v) rewards non-repairs when used as an optimization target, since a compiler-feedback loop raises compile rate while similarity to the human fix falls, with manual inspection finding deletion- and placeholder-style non-repairs among the newly compiling outputs. The natural fallback, whole-function CodeBLEU, also fails: an unchanged copy of the vulnerable input outscores every model. We also examine diff_F1, a change-aware screen that scores only the edited region. It gives exactly zero credit to a no-op and near-zero credit to some, though not all, of the deletion-based gaming patches we observed, while still crediting genuine partial edits, so it may serve as a cheap screen before deeper, execution-based analysis. It is not a repair-quality metric, and we report where it falls short. Our findings argue for change-aware, execution-grounded evaluation of LLM-based vulnerability repair.
发表机构
- The University of Southern Mississippi(南密西西比大学)
机构由 AI 辅助整理,请以论文原文为准。