AI 中文总结
研究修复提交与详细漏洞报告可用性的延迟问题,提出PoCEvolve框架,可从漏洞修复提交直接生成概念验证漏洞利用程序,通过评估相关上下文指导提示进化,在不同模型上有不同成功率提升。
AI 中文摘要
理想情况下,漏洞的详细信息应与修复提交一起提供。但实际上,即便已发布CVE,这些细节往往在提交很久后才出现。此期间,补丁已公开,攻击者可逆向工程,而防御者却缺乏评估暴露情况、确定优先级和验证修复所需的细节。可执行证据(如概念验证漏洞利用程序)能填补这一空白。先前工作已实现漏洞利用程序生成自动化,但最新方法PoCGen假定已有详细漏洞报告,而这正是此期间所缺少的。本文首先进行实证研究,量化修复提交与详细漏洞报告可用性之间的长时间延迟。然后引入PoCEvolve,这是一个漏洞感知提示进化框架,可直接从漏洞修复提交生成漏洞利用程序。给定漏洞修复提交,PoCEvolve会合成相应的漏洞利用程序。为从失败的生成尝试中学习,PoCEvolve评估与漏洞相关上下文不同维度的有用性,包括推断的易受攻击API和代码覆盖信息。这些评估指导提示进化,以生成更有效的漏洞利用程序提示。我们在该http URL上评估PoCEvolve,其漏洞利用程序生成成功率为58.4%,比PoCGen相对提高20.7%,比GPT-4o-mini的大语言模型基线提高200.0%。使用最新模型Qwen3.7-Plus,PoCEvolve成功率更高,达85.3%。当有详细漏洞报告时,PoCEvolve成功率为71.7%,比PoCGen提高11.1%。
英文摘要
Ideally, the detailed information about a vulnerability should be made available together with the fixing commit. In practice, however, such details often become available only long after the commit, even when a CVE has already been published. During this window, the patch is already public, so attackers can reverse-engineer it, yet defenders lack the details needed to assess exposure, prioritize, and validate the fix. Executable evidence, such as a proof-of-concept (PoC) exploit, could fill this gap. Prior work has automated PoC generation, but the state-of-the-art approach, PoCGen, assumes that a detailed vulnerability report is already available, which is precisely what is missing during this window. In this paper, we first present an empirical study quantifying the long delay between the fixing commit and the availability of a detailed vulnerability report. We then introduce PoCEvolve, a vulnerability-aware prompt-evolution framework that generates PoCs directly from vulnerability-fixing commits. Given a vulnerability-fixing commit, PoCEvolve synthesizes a corresponding PoC exploit. To learn from unsuccessful generation attempts, PoCEvolve assesses the usefulness of different dimensions of vulnerability-related context, including the inferred vulnerable API and code-coverage information. These assessments guide prompt evolution towards more effective exploit-generation prompts. We evaluate PoCEvolve on SecBench.VFC.js, where PoCEvolve achieves a PoC generation success rate of 58.4%, corresponding to relative improvements of 20.7% over PoCGen and 200.0% over the LLM baseline with GPT-4o-mini. With a recent model, Qwen3.7-Plus, PoCEvolve achieves a higher success rate of 85.3%. When detailed vulnerability reports are available, PoCEvolve achieves a success rate of 71.7%, improving over PoCGen by 11.1%.