CodeMechanic:基于漏洞属性引导的程序缓解方案
CodeMechanic: Bug-Property-Guided Program Mitigation
浏览论文内容
中文总结 AI 辅助
本文提出CodeMechanic系统,通过重构内存安全属性、插入故障停止保护等方法,在101个ARVO漏洞上生成更多合理且语义等价的补丁,同时降低token使用量。
中文摘要 AI 辅助
自动化测试发现漏洞的速度快于开发人员调查和修复的速度,导致已知内存损坏仍存在可被利用的时间窗口。端到端大语言模型(LLM)修复代理可缩短该时间窗口,但它们会生成无限制的代码变更,且通常仅通过重放概念验证(PoC)来验证变更。这种弱判定标准会接受通过修改无关行为来抑制观察到的崩溃的补丁,使补丁在实际部署中存在风险。本文提出CodeMechanic,这是一个用于生成针对空间内存损坏的受限缓解方案的漏洞属性引导系统。CodeMechanic不要求LLM生成永久修复,而是从崩溃中重构被违反的内存安全属性,验证解引用指针及其缓冲区范围,并在危险访问前插入局部故障停止保护。该保护会在边界检查失败时终止执行,由此产生的缓解方案刻意用可用性换取安全性:可将潜在的远程代码执行转化为受控终止,让开发人员有时间调查根本原因并准备永久修复。CodeMechanic结合二维静态与动态上下文提取器、提示内调试知识及逐步验证,以限制LLM错误的影响。在101个真实世界ARVO漏洞上,CodeMechanic首次尝试生成的合理补丁(即通过PoC重放验证的补丁)比最优基准多47.6%,同时使用的token少91%。人工审计进一步显示,CodeMechanic生成的与开发人员编写的修复在语义上等价的补丁数量是基准的3.4至4.3倍。
英文摘要
Automated testing discovers vulnerabilities faster than developers can investigate and repair them, leaving an interval in which known memory corruptions remain exploitable. End- to-end LLM repair agents can shorten this interval, but they synthesize open-ended code changes and commonly validate them only by replaying a proof of concept (PoC). This weak oracle accepts patches that silence the observed crash by changing unrelated behavior, making unintended deployment risky. We present CodeMechanic, a bug-property-guided system for generating constrained mit- igations for spatial memory corruption. Instead of asking an LLM to generate a permanent repair, CodeMechanic reconstructs the violated memory-safety property from the crash, validates the dereferenced pointer and its buffer range, and inserts a local fail-stop guard before the dangerous access. The guard terminates execution when the boundary check fails. The resulting mitigation deliberately trades availability for security: it can convert potential remote code execution into controlled termination while developers investigate the root cause and prepare a permanent repair. CodeMechanic combines a two-dimensional static and dynamic context extractor with in-prompt debugging knowledge and stepwise val- idation to limit the effect of LLM errors. On 101 real-world ARVO bugs, the first attempt of CodeMechanic produces 47.6% more plausible patches (i.e., patches that pass PoC- replay validation) than the best baseline while using 91% fewer tokens. Manual audit further shows that CodeMechanic produces 3.4x - 4.3x more patches semantically equivalent to developer-written repairs.