GapForge:通过覆盖差距分析进行定向编译器模糊测试
GapForge: Directed Compiler Fuzzing via Coverage-Gap Analysis
AI总结:
研究针对现代编译器代码库覆盖难问题,提出基于大型语言模型的GapForge测试生成技术,通过覆盖驱动评分、路径差异分析等步骤,将覆盖差距作为区域级目标,显著优于现有技术,提高了编译器覆盖率并发现多个实际故障。
AI中文摘要:
现代编译器代码库(如GCC和LLVM)规模大且复杂,全面覆盖不同代码区域极具挑战。现有测试生成技术大多忽略目标代码特性,导致大量编译器区域测试不足。为提高编译器覆盖率,尤其是难以触及的边缘区域,我们提出GapForge,一种基于大型语言模型的定向测试生成技术。它分三步将覆盖差距视为显式区域级目标:首先通过覆盖驱动评分优先处理大的未覆盖文件;其次将每个未覆盖行跨度与其封闭的已覆盖上下文配对并进行路径差异分析以推断细粒度触发要求;最后根据这些要求和先前失败的提示合成提示,利用覆盖反馈指导下一轮选择。在GCC 14.3.0和LLVM 19.1.0上,GapForge显著优于八种先进技术,72小时内分别在GCC和LLVM的核心编译器模块上实现68.13%和69.11%的覆盖率,还发现12个实际编译器故障。
英文摘要:
Modern compiler codebases (e.g., GCC and LLVM) are large and complex, making comprehensive coverage across diverse code regions highly challenging. Most existing test generation techniques ignore characteristics of the target code, producing test programs that exercise only a limited subset of it. Consequently, substantial compiler regions remain insufficiently tested, leaving persistent long-tail coverage gaps that survive across releases. Even existing white-box techniques achieve limited coverage on large-scale compilers. To improve compiler coverage, especially for hard-to-reach edge regions, we present GapForge, a targeted LLM-based test generation technique that reasons about coverage gaps. Unlike program-driven techniques that generate diverse inputs without modeling which regions they exercise, and unlike whole-file summarization that yields coarse guidance, GapForge treats coverage gaps as explicit region-level targets in three steps. First, it prioritizes files via coverage-driven scoring that favors large, undercovered files. Second, it pairs each uncovered line span with its enclosing covered context and performs path-difference analysis to infer fine-grained triggering requirements: the program structures and compilation options needed to reach the uncovered region. Third, it synthesizes prompts from these requirements and previously failed prompts, using coverage feedback to guide next-round selection. On GCC 14.3.0 and LLVM 19.1.0, GapForge significantly outperforms eight state-of-the-art techniques. Within 72 hours, it achieves 68.13% and 69.11% coverage on core compiler modules in GCC and LLVM, surpassing the white-box technique WhiteFox by 24,736 and 19,798 additional lines, respectively. Moreover, GapForge discovers 12 real-world compiler failures (5 in GCC, 7 in LLVM), including 8 crashes and 4 miscompilations, with each component contributing to its performance.