TyPatch:将补丁转化为类型状态规则以检测内核漏洞
TyPatch: Transforming Patches into Typestate Rules for Kernel Bug Detection
浏览论文内容
中文总结 AI 辅助
TyPatch通过LLM将补丁转为类型状态规则,由共享后端执行,解耦缺陷语义与分析器实现,在Linux v6.16上发现559个漏洞,121个获确认,且生成令牌减少88.3%以上,精确度提升至3.42-14.95倍。
中文摘要 AI 辅助
历史Linux内核补丁蕴含的缺陷知识可应用于其原始修复位置之外。近期研究表明,大型语言模型(LLM)能从历史补丁生成静态分析检查器,并用于发现新的内核漏洞。然而,生成完整检查器要求模型既能恢复补丁所表达的缺陷语义,又能实现复杂的程序分析机制,包括对象跟踪、别名分析、路径状态维护和过程间传播。在单一端到端代码生成任务中耦合这些职责,可能将简单的缺陷规则转化为不稳定且昂贵的分析器实现问题。为解决此问题,我们提出TyPatch,将补丁特定的缺陷语义与分析器实现解耦。LLM将每个补丁翻译为类型状态规则,指定其跟踪对象、动作、守卫、转换和违规。共享后端随后执行这些规则,将动作绑定到程序事件,跨别名跟踪对象身份,沿程序路径传播类型状态,并为所有规则生成报告。在Linux v6.16上,TyPatch发现559个不同漏洞,其中121个已获内核开发者确认。在与最先进的完整检查器构建工作流进行匹配的38个补丁对比中,TyPatch使用的生成令牌减少88.3%-90.1%,而其初始报告池的精确度达到该工作流产出结果的3.42-14.95倍。
英文摘要
Historical Linux kernel patches capture defect knowledge that applies beyond their original repair sites. Recent work has shown that large language models (LLMs) can generate static-analysis checkers from historical patches and use them to uncover new kernel bugs. However, complete-checker generation requires the model both to recover the defect semantics expressed by a patch and to implement sophisticated program-analysis machinery, including object tracking, alias analysis, path-state maintenance, and interprocedural propagation. Coupling these responsibilities in a single end-to-end code-generation task can turn a simple defect rule into an unstable and expensive analyzer-implementation problem. To address this problem, we present TyPatch, which decouples patch-specific defect semantics from analyzer implementation. An LLM translates each patch into a typestate rule specifying its tracked object, actions, guards, transitions, and violations. A shared backend then executes these rules, binding their actions to program events, tracking object identity across aliases, propagating typestate along program paths, and producing reports for all rules. On Linux v6.16, TyPatch finds 559 distinct bugs, 121 of which have been confirmed by kernel developers. In a matched 38-patch comparison with the state-of-the-art complete-checker construction workflow, TyPatch uses 88.3-90.1% fewer generation tokens, while its initial report pools achieve 3.42-14.95$\times$ the precision of those produced by that workflow.
发表机构
- Tsinghua University(清华大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。