发表机构
Tulane University; New Jersey Institute of Technology(杜兰大学; 新泽西理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对剥离二进制文件中远程间接控制流预测问题,提出ICFlowNet框架,采用双虚拟中心与多任务图学习技术,构建了可信评估流程,在远程间接调用任务上显著优于现有基线。
AI 中文摘要
恢复间接控制流(ICF)边是二进制安全分析的基础,但现有方法难以处理远程依赖、孤立不同类型的ICF,且常在易受标签噪声和数据泄漏影响的协议下评估。本文提出ICFlowNet,一个用于剥离二进制文件中远程ICF预测的统一框架。ICFlowNet引入候选感知双虚拟中心,即全局代码中心与全局数据中心,以在远程代码和数据证据间创建短路由路径,并将其与多任务图学习结合,联合建模间接调用、间接尾调用、跳转表和返回操作。为实现可信评估,我们进一步开发了一种感知泄漏、可控噪声的评估流程,包含包级拆分、函数级助记符哈希去重,以及由动态正例和绝对负例构建的干净测试协议。利用该流程,我们构建了包含15901个唯一剥离x86-64二进制文件的数据集,其中1351个带有动态真值。实验表明,单纯扩展静态监督仅能带来微小提升,而我们的结构和多任务设计至关重要:双虚拟中心使远程F1值提升最多9.13个百分点,多任务学习带来最多5.81个百分点的提升,最终模型在远程间接调用上的F1值较现有基线高出13个百分点以上,且仅增加11.44%的拓扑开销。
英文摘要
Recovering indirect control-flow (ICF) edges is fundamental to binary security analysis, yet existing methods struggle with long-range dependencies, isolate different ICF types, and are often evaluated under protocols vulnerable to label noise and data leakage. We present ICFlowNet, a unified framework for long-range ICF prediction in stripped binaries. ICFlowNet introduces candidate-aware Dual Virtual Hubs, a Global Code Hub and a Global Data Hub, to create short routing paths between distant code and data evidence, and combines them with multi-task graph learning to jointly model indirect calls, indirect tail calls, jump tables, and returns. To enable credible evaluation, we further develop a leakage-aware, noise-controlled pipeline with package-level splits, function-level mnemonic-hash deduplication, and a clean test protocol built from dynamic positives and absolute negatives. Using this pipeline, we construct a dataset of 15,901 unique stripped x86-64 binaries, including 1,351 with dynamic ground truth. Experiments show that simply scaling static supervision yields only marginal gains, whereas our structural and multi-task designs are essential: Dual Virtual Hubs improve long-range F1 by up to 9.13 points, multi-task learning adds up to 5.81 points, and the final model outperforms prior baselines by more than 13 F1 points on long-range indirect calls while adding only 11.44 percent topological overhead.
CommentsAccepted by ACM CCS 2026