arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12457cs.SE

Hieronym:利用分层多源信息进行剥离二进制函数重命名

Hieronym: Leveraging Hierarchical Multi-Source Information for Function Renaming in Stripped Binary

  • Zhongguancun Laboratory(中关村实验室)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaoling Zhang, Jian Sun, Dawei Wang, Chongyu Wang, Li Chen, Zhaoteng Yan, Peipei Liu, Lixiao Zhang, Dan Li

中文总结 AI 辅助

Hieronym提出一种基于大语言模型的分层摘要驱动框架,整合全局与局部上下文及函数语义,在四种架构和优化级别下显著提升剥离二进制函数重命名性能,并具强泛化能力。

中文摘要 AI 辅助

剥离二进制中的函数重命名可以通过提高代码可读性来极大地辅助逆向工程师,然而这是一项具有挑战性的任务。其难度源于需要从跨不同指令集、架构和编译器优化的低级二进制代码中准确捕获函数语义,并以简洁、人类可读的名称表达这些语义。现有方法要么无法充分捕获全面的函数语义,要么对未见过的二进制文件表现出有限的泛化能力。在本文中,我们提出了Hieronym,一个基于生成式大语言模型(LLM)的剥离二进制函数重命名框架。Hieronym采用了一种分层摘要驱动的领域自适应策略,并整合了多源信息,包括全局二进制上下文、局部调用上下文和内在函数语义,以增强LLM对二进制代码的理解。为了进行系统评估,我们进一步提出了一个双层评估框架,该框架结合了令牌级和全名级指标。我们在四种架构(x64、x86、ARM和MIPS)下以四种编译器优化级别(O0-O3)编译的二进制函数上评估了Hieronym。实验结果表明,Hieronym显著优于最先进的方法,在精确率、召回率和F1分数上分别实现了50.12%、41.75%和45.10%的令牌级提升,以及79.94%的名称级准确率提升,同时还表现出强大的泛化能力。此外,对真实世界恶意软件样本的实验进一步验证了Hieronym在安全关键场景中的实际有效性。

英文摘要

Function renaming in stripped binaries can substantially assist reverse engineers by improving code readability, yet it is a challenging task. The difficulty stems from the need to accurately capture function semantics from low-level binary code across diverse instruction sets, architectures, and compiler optimizations, and to express these semantics in concise, human-readable names. Existing approaches either inadequately capture comprehensive function semantics or exhibit limited generalization to previously unseen binaries. In this paper, we present Hieronym, a generative large language model (LLM)-based framework for stripped binary function renaming. Hieronym adopts a hierarchical summarization-driven domain adaptation strategy and integrates multi-source information, including global binary context, local calling context, and intrinsic function semantics, to enhance the LLM's understanding of binary code. To enable systematic evaluation, we further propose a dual-layer evaluation framework that incorporates both token-level and whole-name-level metrics. We evaluate Hieronym on binary functions compiled with four compiler optimization levels (O0-O3) for four architectures (x64, x86, ARM, and MIPS). Experimental results demonstrate that Hieronym significantly outperforms state-of-the-art methods, achieving token-level improvements of 50.12% in precision, 41.75% in recall, and 45.10% in F1-score, as well as a 79.94% improvement in name-level accuracy, while also exhibiting strong generalization capability. Moreover, experiments on real-world malware samples further validate the practical effectiveness of Hieronym in security-critical scenarios.

补充信息

↑