发表机构
University of Massachusetts Amherst; Northwestern University(马萨诸塞大学阿默斯特分校; 西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出鲁棒性评估框架和分裂(splitting)方法,通过复制模糊器队列状态分支执行,在Magma和FuzzBench上显著提升漏洞发现率并降低变异性。
AI 中文摘要
基于突变的模糊测试被广泛用于发现软件漏洞,但其随机性使得严格评估和可靠漏洞检测变得复杂。先前的工作通过经验方式衡量这种变异性,但缺乏具有可计算收敛性和样本复杂度保证的理论。我们同时解决这两个问题。首先,我们在考虑计算量后,通过衡量漏洞触发率的变化,从独立测试活动中估计鲁棒性。对于长度为 $T$ 的 $M$ 次测试活动,有限试验误差以标准的 $M^{-1/2}$ 速率下降,而当时间上的漏洞相关性衰减时,有限长度误差有界。随后,我们引入分裂(splitting),这是一种黑盒包装器,在漏洞触发后复制模糊器的队列状态,并从该状态在多个分支中继续,将更多努力导向已发现的区域。对于任何实现的分裂树,相对于单一继续路径,分支不会减少漏洞触发事件的原始数量。当接收更多分支的状态倾向于在后续产生更多漏洞时,预期检测也会改善。在一个简化模型中,当 $p<\sqrt{2}-1$ 时,分裂降低了每单位计算的方差,其中 $p$ 是花费在漏洞区域的时间比例。在匹配计算量下,分裂在40个Magma基准真实单元中的38个中,每CPU小时发现的真实漏洞多于基线(中位数+52%;38个中有34个单独显著),并且从未发现更少的独特漏洞。它将CVE-2019-19926从未检测(0/20次试验)变为可靠检测(20/20;Fisher $p<10^{-4}$),并有六项额外的检测改进,其中五项涉及CVE。在FuzzBench上,分裂在70对中的53对中发现了更多独特漏洞,并在70对中的66对中减少了跨测试活动的变异性,中位数减少约$10\times$。每个分支都计入计算预算。以约0.14%的开销,分裂提供了一种实用方法来衡量和改进模糊器。
英文摘要
Mutation-based fuzzing is widely used to discover software vulnerabilities, but its randomness complicates rigorous evaluation and reliable bug detection. Prior work measures this variability empirically but lacks a theory with computable convergence and sample-complexity guarantees. We address both problems. First, we estimate robustness from independent campaigns by measuring variation in bug-trigger rates after accounting for compute. For $M$ campaigns of length $T$, finite-trial error decreases at the standard $M^{-1/2}$ rate, while finite-length error is bounded when temporal bug correlations decay. We then introduce splitting, a black-box wrapper that copies a fuzzer's queue state after a bug trigger and continues from that state in multiple branches, directing more effort toward the discovered region. For any realized split tree, branching cannot reduce the raw number of bug-triggering events relative to a single continuation path. Expected detection also improves when states receiving more branches tend to yield more bugs later. In a simplified model, splitting reduces variance per unit compute when $p<\sqrt{2}-1$, where $p$ is the fraction of time spent in the bug region. At matched compute, splitting finds more real bugs per CPU-hour than the baseline in 38 of 40 Magma ground-truth cells (median +52\%; 34 of 38 individually significant) and never finds fewer distinct bugs. It changes CVE-2019-19926 from undetected (0/20 trials) to reliably detected (20/20; Fisher $p<10^{-4}$), with six additional detection improvements, five involving CVEs. On FuzzBench, splitting finds more unique bugs in 53 of 70 pairs and reduces cross-campaign variation in 66 of 70, with a median reduction of about $10\times$. Every branch counts toward the compute budget. With about 0.14\% overhead, splitting provides a practical way to measure and improve fuzzers.