发表机构
School of Computer Science and Engineering, Nanjing University of Science and Technology; Department of Computer Science, Xidian University; School of Computer Science, Wuhan University(南京理工大学计算机科学与工程学院; 西安电子科技大学计算机科学与技术学院; 武汉大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Reprise通过提炼已知缺陷为语义图模式并在图合成中规避这些模式,在DL编译器模糊测试中有效抑制已知缺陷报告,同时保持覆盖率并发现新缺陷。
AI 中文摘要
模糊测试在发现深度学习(DL)编译器中的缺陷方面非常有效,但现有的模糊测试器会反复触发它们已经发现的故障。当前的模糊测试器通过多样化生成的程序来工作,下游工具则事后对缺陷报告进行去重,但这两者都无法阻止模糊测试器生成会重新触发已知缺陷的程序。我们提出了Reprise,一种在测试生成过程中抑制已知缺陷报告的DL编译器模糊测试器。Reprise采用了一个刻意轻量级的生成器,在统一中间表示(UIR)中构建图,并带有局部算子签名。与多样化程序不同,Reprise将每个发现的缺陷提炼成一个语义图模式,该模式捕获了触发该缺陷所需的算子、值约束、图上下文和数据流。在图合成过程中,它会重新生成任何完成已知模式的节点,从而在编译和执行之前避免已知缺陷的触发。我们在三个DL编译器上评估了Reprise:TVM、PyTorch Inductor和ONNX Runtime。与无引导的变体相比,针对主要已知缺陷编写的模式在TVM上将崩溃报告减少了90.2%(从254降至25),在PyTorch Inductor上减少了93.8%(从594降至37),同时保持了可比较的分支覆盖率,并且执行的测试减少了1.0%-17.9%。这些结果来自每种配置的一次运行,并且涉及那些模式所提炼的缺陷。Reprise还在三个编译器中发现了25个以前未知的缺陷,其中24个在一个月内被发现,其中7个已被修复,另外12个已得到开发者的确认。复制包已在[11]中提供。
英文摘要
Fuzzing is effective at finding bugs in deep learning (DL) compilers, but existing fuzzers repeatedly trigger faults they have already uncovered. Current fuzzers diversify the generated programs, and downstream tools deduplicate bug reports post hoc, but neither stops a fuzzer from generating programs that re-trigger known defects. We present Reprise, a DL compiler fuzzer that suppresses reports of known defects during test generation. Reprise uses a deliberately lightweight generator that builds graphs in a unified intermediate representation (UIR) with local operator signatures. Instead of diversifying programs, Reprise distills each discovered defect into a semantic graph pattern that captures the operators, value constraints, graph context, and data flow required to trigger it. During graph synthesis, it regenerates any node that completes a known pattern, so known defect triggers are avoided before compilation and execution. We evaluate Reprise on three DL compilers: TVM, PyTorch Inductor, and ONNX Runtime. Against its unguided variant, patterns written for the dominant known defects reduce crash reports by 90.2% on TVM (254 to 25) and by 93.8% on PyTorch Inductor (594 to 37), with comparable branch coverage and 1.0-17.9% fewer executed tests. These results come from one run per configuration and concern the defects the patterns were distilled from. Reprise also found 25 previously unknown bugs across the three compilers, 24 of them within one month, of which 7 have been fixed and 12 more confirmed by developers. The replication package has been provided at [11].