发表机构
Shandong University; Hong Kong University of Science and Technology(山东大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对二进制代码表示学习忽略指令级对应关系的问题,提出将指令对齐作为辅助训练目标,经实验验证可提升检索精度与相似性判断的区分度。
AI 中文摘要
二进制代码表示学习是软件安全与逆向工程领域的基础问题。现有方法主要学习函数级嵌入,捕捉二进制函数间的粗粒度语义关系,但很大程度上忽略了细粒度的指令级对应关系。这一局限使得模型无法利用编译器调试信息提供的宝贵监督信号,而这类信号本可助力学习更准确、可解释的二进制代码表示。我们提出利用指令对齐知识进一步改进二进制代码表示学习。初步研究显示,针对函数级二进制代码相似性微调后的模型,其指令对齐表现远优于预训练模型,表明指令对齐与函数级嵌入质量存在强相关性。基于此,我们设计了一种训练方法,将指令对齐明确纳入辅助训练目标。实验表明,指令对齐训练可提升检索精度,并为模型的相似性判断提供更具区分度的信号。
英文摘要
Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn function-level embeddings that capture coarse-grained semantic relationships between binary functions, but they largely ignore fine-grained instruction-level correspondences. This limitation misses valuable supervision signals available from compiler debug information, which can support the learning of more accurate and interpretable binary code representations. We propose to leverage instruction alignment knowledge to further improve binary code representation learning. Our preliminary study reveals that models finetuned for function-level binary code similarity exhibit substantially better instruction alignment than their pre-trained model, suggesting a strong correlation between instruction alignment and function-level embedding quality. Motivated by this observation, we design a training approach that explicitly incorporates instruction alignment as an auxiliary training objective. Our experiments show that instruction alignment training improves retrieval accuracy and provides more discriminative signal for the model's similarity judgments.
CommentsIn proceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)