arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不透明指针下的神经符号间接调用分析

Neuro-Symbolic Indirect-Call Analysis under Opaque Pointers

Kaixuan Li, Bozhi Wu, Jian Zhang, Peixin Wang, Ting Su, Yang Liu

arXiv 2609.33547首次发表:更新:

发表机构

Nanyang Technological University; Beihang University; East China Normal University(南洋理工大学; 北京航空航天大学; 华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对不透明指针下C语言间接调用解析难题,提出Facet分析,重建字段分派关系,结合LLM处理残余情况,显著提升调用图精度并发现深层错误。

AI 中文摘要

解析间接调用是构建 C 语言调用图的核心问题。诸如 MLTA 等可扩展的基于类型的分析利用 LLVM IR 中的类型信息,将间接调用与赋值给相应结构体字段的函数关联起来。然而,单个指向类型常常无法准确表示指针所指向的内存,并且 LLVM 17 移除了指向类型,转而采用不透明指针。因此,字段敏感分析失去了其匹配键。恢复被擦除的类型可以重新获得匹配键,但仍会遗漏类型所编码的关系:程序将哪些函数赋值给该字段。我们提出了 Facet,据我们所知,这是第一个在不透明 IR 上重建这种分派关系的分析。Facet 识别出间接调用加载其函数指针所来自的结构体字段。它分别通过初始化器、存储和聚合复制来恢复赋值给该字段的函数。然后,它通过字段标识将两者连接起来,无需端到端的值流路径。Facet 在边添加和移除的不同证据规则下对提议的调用图变更进行分类,并记录每次细化背后的假设。一个 LLM 仅决定符号有界候选者中的剩余情况。一种分析同时产生保留召回率的调用图和细化后的调用图。在 14 个 C 程序上,Facet 将平均目标集大小从 25.9 降至 5.2,并将观察到的召回率从 0.79 提升至 0.99。其恢复的字段标识在 98.1% 的联合解析位点与类型化 IR 一致。应用于错误检测时,细化后的调用图在从 nginx 到 Linux 内核的 C 软件中发现了 17 个深层错误,其中三个潜伏了十多年;12 个已确认。

英文摘要

Resolving indirect calls is central to call-graph construction for C. Scalable type-based analyses such as MLTA use type information in LLVM IR to associate indirect calls with functions assigned to the corresponding structure fields. However, a single pointee type often misrepresents the memory a pointer addresses, and LLVM 17 removed pointee types in favor of opaque pointers. Therefore, field-sensitive analyses lose their matching key. Recovering the erased types restores the matching key but still misses the relation that the type encoded: which functions the program assigns to the field. We present Facet, to our knowledge the first analysis that reconstructs this dispatch relation over opaque IR. Facet identifies the structure field from which an indirect call loads its function pointer. It separately recovers the functions assigned to that field through initializers, stores, and aggregate copies. It then joins the two by field identity, without requiring an end-to-end value-flow path. Facet classifies proposed call-graph changes under distinct evidence rules for edge addition and removal and records the assumption behind each refinement. An LLM decides only the residual cases among symbolically bounded candidates. One analysis yields both a recall-preserving call graph and a refined call graph. On 14 C programs, Facet reduces the mean target-set size from 25.9 to 5.2 and raises observed recall from 0.79 to 0.99. Its recovered field identities agree with typed IR at 98.1% of jointly resolved sites. Applied to bug detection, the refined call graph found 17 deep bugs in C software from nginx to the Linux kernel, three of them latent for over a decade; 12 are confirmed.

CommentsRevised version with corrected formatting

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑