COMMITGUARD:用于检测提交诱导型漏洞的差分切片模糊测试方法
COMMITGUARD: Differential Slice Fuzzing for Commit-Induced Bug Detection
浏览论文内容
中文总结 AI 辅助
COMMITGUARD是一种基于提交感知差分切片的模糊测试方法,通过对比提交前后代码切片的 sanitizer 报告,从300次开源项目提交中识别出7份候选漏洞,其中5份为真实漏洞,平均分析单提交耗时32.4分钟,修改函数覆盖率达75.36%。
中文摘要 AI 辅助
现代软件系统通过频繁的提交来实现漏洞修复、功能添加和安全补丁。尽管代码审查和测试被广泛用于检查这些变更,但它们对内存安全问题的保证往往有限。代码审查人员可能会遗漏微妙的边界、生命周期或初始化错误,而现有测试可能无法覆盖提交所影响的特定路径。模糊测试在暴露此类漏洞方面很有效,但将其应用于每一次提交仍不切实际,因为全程序模糊测试成本高昂,需要合适的测试 harness,且仍可能无法触及提交所变更的代码。本文中,我们提出了COMMITGUARD,一种基于提交感知差分切片的模糊测试方法,用于验证代码变更。COMMITGUARD背后的核心见解是,被修改函数的提交前版本可作为解释提交后发现漏洞的行为基线。对于每个目标提交,COMMITGUARD会识别被修改的函数,从提交前和提交后版本中提取可编译的代码切片,并对成对的切片独立进行模糊测试。随后,它会比较两个版本的 sanitizer 报告,并将仅在提交后版本中出现的漏洞报告为候选提交诱导型漏洞。我们在来自openSSL、libpcap和leptonica的300次提交上评估了COMMITGUARD。切片模糊测试最初在这些提交上产生了518份 sanitizer 报告。通过比较提交前和提交后的切片,COMMITGUARD将这一大量输出缩小到7份需要手动分类的候选提交诱导型漏洞报告。手动验证确认其中5份报告为真实漏洞,在我们报告后,这些漏洞已被所检查项目的开发人员修复,而仅有2份报告被归类为误报。COMMITGUARD分析一次提交平均耗时32.4分钟,对被修改函数的平均覆盖率达到75.36%。
英文摘要
Modern software systems evolve through frequent commits that implement bug fixes, features, and security patches. Although code review and testing are widely used to check these changes, they often provide limited assurance for memory-safety issues. Code reviewers may miss subtle boundary, lifetime, or initialization errors, while existing tests may not exercise the specific paths affected by a commit. Fuzzing is effective at exposing such bugs, but applying it to every commit remains impractical because whole-program fuzzing is expensive, requires suitable harnesses, and may still fail to reach the code changed by a commit. In this paper, we introduce COMMITGUARD, a commit-aware differential slice-based fuzzing approach for verifying code changes. The key insight behind COMMITGUARD is that the pre-commit version of a modified function can serve as a behavioral baseline for interpreting bugs found after the commit. For each target commit, COMMITGUARD identifies modified functions, extracts compilable code slices from both the pre-commit and post-commit versions, and fuzzes the paired slices independently. It then compares sanitizer reports across the two versions and reports bugs that emerge only in the post-commit version as candidate commit-induced bugs. We evaluate COMMITGUARD on 300 commits from openSSL, libpcap and leptonica. Slice fuzzing initially produces 518 sanitizer reports across these commits. By comparing pre-commit and post-commit slices, COMMITGUARD narrows this large output to 7 candidate commit-induced bug reports that require manual triage. Manual validation confirms 5 of these reports as real bugs that were fixed by developers of the examined projects after we reported them, while only 2 reports were classified as false positives. COMMITGUARD analyzes a commit in 32.4 minutes on average and achieves 75.36% average coverage of modified functions.