自动化安全补丁回移植基准测试:我们还差多远?
Benchmarking Automated Security Patch Backporting: How Far Are We?
浏览论文内容
中文总结 AI 辅助
该研究推出涵盖多场景的Porting Benchmark基准,评估五类安全补丁回移植工具,发现工具泛化能力差、复杂补丁性能骤降,明确根本原因并指出优化方向。
中文摘要 AI 辅助
自动化安全补丁回移植对于缓解N-day漏洞至关重要。现有工具在各自数据集上报告的成功率超过80%,但这些评估往往局限于同构环境,如单一代码库或特定项目版本,因此这些工具在超出原始目标场景时的泛化能力仍不明确。我们推出Porting Benchmark,这是一个包含1234个安全补丁回移植案例的精选数据集,涵盖跨版本、跨分支和跨代码库场景,并配套通用评估框架。我们使用该基准在对齐设置下评估了五种工具,涵盖程序分析、LLM提示和LLM智能体。结果显示,对齐评估改变了表面性能格局:PortGPT和TSBPort在复制数据集上仍相对较强,而FixMorph和Mystique在通用协议下性能大幅下降。在结构复杂的补丁上性能急剧下降:最佳提交级成功率从I型补丁的85.2%降至IV型的24.0%。我们确定了四类根本原因(缺失目标API感知、跨版本语义不匹配、非本地依赖传播失败以及补丁构建或定位失败),并推导了下一代工具设计的具体方向。在包含验证测试用例和构造POC的45个动态验证子集上,我们进一步观察到,基于参考的基准评分无法完全捕捉现实世界的修复效果:精确匹配严重低估了更困难的目标适配,而可执行验证揭示了静态参考一致性遗漏的目标中残留的集成失败,可执行反馈优化在最困难的可执行案例上提供了有限但可衡量的恢复效果。
英文摘要
Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective datasets. However, these evaluations are often confined to homogeneous environments, such as one repository or specific project versions. Consequently, it remains unclear how well these tools generalize beyond their originally targeted scenarios. We present Porting Benchmark, a curated dataset of 1,234 security patch backporting cases spanning cross-version, cross-branch, and cross-repository scenarios, paired with a common evaluation framework. Using this benchmark, we evaluate five tools spanning program analysis, LLM prompting, and LLM agents under aligned settings. Our results show that aligned evaluation changes the apparent performance landscape: PortGPT and TSBPort remain comparatively strong on the Replication Dataset, while FixMorph and Mystique degrade substantially under the common protocol. Performance degrades sharply on structurally complex patches: the best commit-level success rate falls from 85.2% on Type-I patches to 24.0% on Type-IV. We identify four root-cause categories (missing target API awareness, cross-version semantic mismatch, non-local dependency propagation failure, and patch construction or localization failure) and derive concrete directions for next-generation tool design. On a 45-case dynamically validated subset with verified test cases and constructed POCs, we further observe that reference-based benchmark scores do not fully capture real-world remediation: exact match sharply under-credits harder target adaptations, while executable validation reveals residual integration failures in the target that static reference agreement misses. Executable-feedback refinement provides limited but measurable recovery on the hardest executable cases.
发表机构
- Xidian University(西安电子科技大学)
- Nankai University(南开大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。