AI 中文总结
本研究提出Vul4Py基准测试,含100个Python真实漏洞及配对预言机,对比三类AVR方法,发现智能体OpenHands修复漏洞数远超同类方法,配对预言机保障结果可信。
AI 中文摘要
自动化漏洞修复(AVR)在程序分析、机器学习及大语言模型(LLM)领域已取得快速进展,但目前仍缺乏针对Python的可验证、直接对比的AVR方法比较。Python是关键Web、数据及机器学习基础设施的核心支撑,然而现有Python基准测试仅通过概念验证漏洞利用就认可补丁,或仅对上游项目恰好提供功能测试的条目子集应用功能测试,两者均遗漏了功能回归——即补丁虽能挫败漏洞利用,却破坏了无关行为。我们提出Vul4Py,这是一个Python AVR基准测试,其中每个条目都带有配对预言机:一个漏洞利用预言机,在易受攻击版本上必须失败、在修复版本上必须通过;以及一个项目原生的pytest功能预言机,在两个版本上都必须通过。Vul4Py包含来自60个开源项目的100个真实漏洞,涵盖60种不同的CWE(常见弱点枚举),时间跨度为2017年至2025年,每个漏洞都附带固定的、可复现的实例化环境。利用Vul4Py,我们对三类共六种方法进行了比较:专用漏洞修复工具、直接提示的LLM、软件工程智能体。智能体表现突出:OpenHands修复了100个漏洞中的41个,而最强的直接提示LLM修复了4个,专用工具修复了2个,尽管三者使用相同的主干模型。配对预言机是这些计数可信的原因:它拒绝了仅漏洞利用预言机会接受的119个补丁中的15个,且其认可的104个补丁中有98个经人工确认与开发者的补丁语义等价。
英文摘要
Automated Vulnerability Repair (AVR) has advanced rapidly across program analysis, machine learning, and Large Language Models (LLMs), but a verifiable, head-to-head comparison of AVR approaches on Python is still missing. Python underpins critical web, data, and machine-learning infrastructure, yet existing Python benchmarks accept a patch on the strength of a proof-of-concept exploit alone, or apply a functional test only on the subset of entries whose upstream project happens to ship one. Both therefore miss functional regressions, in which a patch defeats the exploit but breaks unrelated behavior. We present Vul4Py, a Python AVR benchmark in which every entry carries a paired oracle: an exploit oracle that must fail on the vulnerable revision and pass on the fixed one, together with a project-native pytest functional oracle that must pass on both. Vul4Py comprises 100 real vulnerabilities from 60 open-source projects, spanning 60 distinct CWEs and the years 2017 to 2025, each packaged with a pinned, reproducible per-instance environment. Using Vul4Py, we compare six approaches in three categories: a specialized vulnerability repair tool, directly prompted LLMs, and software engineering agents. The agents dominate: OpenHands repairs 41 of 100 vulnerabilities, against 4 for the strongest directly prompted LLM and 2 for the specialized tool, despite all three sharing the same backbone model. The paired oracle is what makes these counts trustworthy: it rejects 15 of the 119 patches that an exploit-only oracle would accept, and 98 of the 104 patches it admits are manually confirmed to be semantically equivalent to the developer's patches