发表机构
University of Massachusetts Amherst; Georgia Institute of Technology(马萨诸塞大学阿默斯特分校; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CAROL利用模糊器当前上下文(奖励趋势、等待时间等)进行在线调度,在Magma基准上相比集成基线多触发漏洞,并在五个C++程序中发现120个新崩溃缺陷。
AI 中文摘要
集成模糊测试(Ensemble fuzzing)会运行多个模糊器(fuzzer),并由一个调度器在它们之间分配CPU时间。现有的调度器基于过去性能的简洁摘要以及活动开始前设定的规则来做出这些决策。我们的测量揭示了两个局限性。首先,基于过去奖励的摘要无法可靠地捕捉性能演化:在考虑估计噪声后,连续窗口排名之间的一致性在统计上与窗口内自身一致性无法区分。其次,预测信号因目标而异:在九个目标中的八个上,从其他八个目标学习到的权重对奖励的预测效果不如在当前目标上学习到的权重。我们提出了CAROL,一种在线调度器,它利用每个模糊器的当前上下文。这些上下文信息已可用于调度循环,描述了奖励趋势、等待时间和平台期时间、已到达的代码以及估计不确定性。CAROL以两种方式使用上下文:一种领域引导的方法检测模糊器是处于上升期还是衰退期,并应用特定阶段的更新规则;而一种学习方法则从15个上下文信号预测奖励,并利用预测不确定性进行在线选择。在九个Magma目标上,每当结果存在差异时,CAROL触发的独特漏洞数量都多于三种集成调度基线中的每一种,且从未少于任何基线。与每个目标的最强基线相比,CAROL提升了11.8%,并超越了那种事后为每个目标选择最佳单一模糊器的理想化方法。移除上下文会消除这种增益,且额外发现的漏洞主要集中在那些基线很少或从未触发的漏洞上。在五个广泛使用的C++程序上不做修改地运行,CAROL发现了120个先前未知的崩溃缺陷,这些缺陷已按位置、故障和入口点去重;所有缺陷均已通过项目声明的披露渠道报告给了维护者。
英文摘要
Ensemble fuzzing runs multiple fuzzers on a target while a scheduler allocates CPU time among them. Existing schedulers base these decisions on compact summaries of past performance and rules fixed before a campaign. Our measurements reveal two limitations. First, past-reward summaries do not reliably capture performance evolution: after accounting for estimation noise, agreement between consecutive-window rankings is statistically indistinguishable from within-window self-agreement. Second, predictive signals vary across targets: on eight of nine targets, a weighting learned from the other eight predicts reward worse than one learned on the current target. We introduce CAROL, an online scheduler that uses each fuzzer's current context. Already available to the dispatch loop, this context describes reward trends, waiting and plateau time, reached code, and estimation uncertainty. CAROL uses context in two ways: a domain-guided method detects whether a fuzzer is rising or rotting and applies a phase-specific learning rule, while a learned method predicts reward from 15 context signals and uses predictive uncertainty for online selection. Across nine Magma targets, CAROL triggers more unique bugs than each of three ensemble-scheduling baselines whenever their results differ, and fewer on none. Compared with the strongest baseline for each target, CAROL gains 11.8% and surpasses an oracle that retrospectively selects the best single fuzzer per target. Removing context eliminates the gain, and the additional bugs are concentrated among those the baselines trigger rarely or never. Run unchanged on five widely used C++ programs, CAROL finds 120 previously unknown crashing defects, deduplicated by site, fault, and entry point; all were reported to maintainers through the projects' stated disclosure channels.