arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18457cs.CRcs.SE

AIJon: 自动化生成模糊测试注释

AIJon: Automated Generation of Annotations for Fuzzing

Jayakrishna Menon Vadayath, Hulin Wang, Moritz Schloegel, Jie Hu, Wil Gibbs, Tiffany Bao, Adam Doupé, Ruoyu "Fish" Wang, Yan Shoshitaishvili

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AIJON系统,利用大型语言模型自动生成IJON风格的注释以指导模糊测试,在Magma基准上评估发现其与人类注释效果相当,但未严格优于AFL++,并分析了注释对模糊测试能量分布的影响。

中文摘要 AI 辅助

现代模糊测试器使用代码覆盖率作为反馈来引导其探索,这已被证明是驱动探索的有效策略。然而,这种策略忽略了那些即使没有发现新代码路径也可能对目标程序有意义的输入。幸运的是,先前的研究表明,由人类领域专家生成的注释可以提供额外的反馈,将模糊测试器引导至程序中有趣的部分。在本文中,我们复现了IJON中提出的实验,并将其扩展到大规模的真实世界漏洞检测。为了缓解因需要人类领域专业知识而带来的可扩展性挑战,我们提出利用大型语言模型(LLM)来自动生成注释。我们证明了LLM在此目的上的适用性,并观察到LLM生成的注释可以与人类生成的注释表现相当。受这一发现的启发,我们设计了AIJON系统,该系统利用LLM自动生成IJON风格的注释。我们在Magma基准上评估了AIJON,并惊讶地观察到基于注释的模糊测试并不严格优于AFL++。我们进行了多项实验以确定我们结果的原因,并识别出关于注释对模糊测试活动影响的关键见解,包括它们对模糊测试器能量分布的影响。值得注意的是,我们观察到LLM可以生成与人类生成的注释取得相当结果的注释,从而为未来研究大规模注释影响打开了大门。

英文摘要

Modern fuzzers use code coverage as feedback to guide their exploration which has proven to be an effective strategy for driving exploration. However, this strategy overlooks inputs that may be interesting to the target program even without uncovering new code paths. Fortunately, prior research has shown that annotations generated by human domain experts can provide additional feedback, guiding the fuzzer towards interesting parts of the program. In this paper, we replicate experiments presented in IJON and extend them to real-world vulnerability detection at scale. To mitigate the scalability challenge, imposed by the need for human domain expertise, we propose utilizing LLMs to automatically generate annotations. We demonstrate the applicability of LLMs for this purpose and observe that LLMs can generate annotations that perform comparably to human-generated annotations. Motivated by this finding, we design AIJON, a system that leverages LLMs to automatically generate IJON-style annotations. We evaluate AIJON on the Magma benchmark and surprisingly observe that annotation-based fuzzing does not perform strictly better than AFL++. We conduct several experiments to identify the cause of our results and identify key insights regarding the impact of annotations on fuzzing campaigns, including their effect on the energy distribution of the fuzzer. Notably, we observe that LLMs can generate annotations that achieve comparable results to human generated ones, thus opening the door for future research to perform further studies on the impact of annotations at scale.

发表机构

  • Arizona State University(亚利桑那州立大学)
  • CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)

机构由 AI 辅助整理,请以论文原文为准。

↑