AI 中文总结
针对现有PDF阅读器模糊测试工具的局限,提出基于LLM的PDFuzzer,生成复杂API调用序列,在三款主流阅读器上覆盖率提升48%,识别31个零日漏洞并获厂商漏洞赏金。
AI 中文摘要
现有PDF阅读器模糊测试工具依赖仅包含单个API调用的简单测试用例,导致代码覆盖率有限,可能遗漏需要序列API调用的漏洞。为解决这些局限,我们提出PDFuzzer,一种新型PDF引擎模糊测试工具,可自动生成复杂且有意义的API调用序列。PDFuzzer首先利用大语言模型(LLM)从JavaScript API手册提取的规范和执行跟踪中构建上下文无关文法,并推断单个API调用间的关系;基于这些文法和关系,PDFuzzer采用约束求解器生成用于模糊测试的具体API调用序列。实验表明,在Adobe Acrobat Reader、Foxit PDF Reader和PDF-XChange Editor三款主流PDF阅读器上,PDFuzzer的性能显著优于当前最先进的PDF模糊测试工具(TypeOracle、Favocado、Cooper)及基于LLM的模糊测试工具(Fuzz4All、 naive LLM);其覆盖率较现有工具最高提升48%,并在上述阅读器中识别出31个零日漏洞,涵盖信息泄露到任意代码执行。我们的 ablation study 验证了各组件的必要性,包括LLM在全流程各阶段的准确率达93%-98%;我们通过协调漏洞披露流程向厂商公开了所有漏洞,并获得了漏洞赏金。
英文摘要
Existing fuzzers for PDF readers rely on simple test cases that involve only individual API calls, leading to limited coverage and potentially missing vulnerabilities that require sequences of API calls. To address these limitations, we propose PDFuzzer, a novel PDF engine fuzzer that automatically generates complex and meaningful API call sequences. PDFuzzer first uses a Large Language Model (LLM) to construct context-free grammars and infer the relationships between individual API calls from specifications extracted from JavaScript API manuals and execution traces. Based on the grammars and relationships, PDFuzzer employs a constraint solver to generate concrete API call sequences for fuzzing. Our experiments show that PDFuzzer significantly outperforms state-of-the-art PDF fuzzers (TypeOracle, Favocado, and Cooper) and LLM-based fuzzers (Fuzz4All, naive LLM) on three mainstream PDF readers: Adobe Acrobat Reader, Foxit PDF Reader, and PDF-XChange Editor. PDFuzzer achieves up to 48% higher coverage than existing tools and identifies 31 zero-day vulnerabilities in these readers, from information leakage to arbitrary code execution. Our ablation study validates the necessity of each component, including LLMs, which achieve high accuracy across all pipeline stages (93-98%). We disclosed all vulnerabilities to the vendors via a coordinated vulnerability disclosure process and received bug bounties.
Comments16 pages, 2 figures, accepted by ACM CCS 2026