arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型的后置注意力引导:面向混淆代码的鲁棒代码理解

Post-Hoc Attention Steering of Large Language Models for Robust Code Understanding under Obfuscation

Xiaokai Rong, Aashish Yadavally, Tien N. Nguyen

arXiv 2609.26102首次发表:更新:

发表机构

University of Texas at Dallas; University of Central Florida(德克萨斯大学达拉斯分校; 中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM在混淆代码上性能下降的问题,提出CodeSteer注意力引导方法,通过重分配注意力至语义相关元素,显著提升混淆代码理解性能,并展示其在漏洞检测中的潜力。

AI 中文摘要

代码混淆被广泛应用于软件系统和恶意软件中,用以隐藏程序逻辑并阻碍分析,这给人类开发者和自动化工具都带来了重大挑战。尽管大语言模型(LLMs)在代码理解方面展现出强大能力,但其对混淆的鲁棒性仍知之甚少。我们的初步研究表明,LLM在混淆代码上的性能显著下降,这表明其依赖表面词汇线索而非深层语义推理。为解决这一局限,我们提出了CodeSteer,一种新颖的注意力引导方法,该方法将模型注意力重新分配到语义相关的程序元素上,包括用于输出预测的后向切片和用于执行推理的控制流路径。我们的方法将轻量级程序分析与推理时注意力引导相结合,以引导LLM关注程序的核心输入到输出依赖关系。跨多个模型和数据集的实验表明,CodeSteer显著提升了在混淆代码上的性能,通常能恢复到与未混淆程序相当的水平。我们还通过一个缓冲区溢出检测的案例研究展示了CodeSteer的实际效用,凸显了其在恶意软件/漏洞分析和混淆代码逆向工程中的潜力。

英文摘要

Code obfuscation is widely used in software systems and malware to conceal program logic and hinder analysis, posing significant challenges for both human developers and automated tools. While large language models (LLMs) have shown strong capabilities in code understanding, their robustness to obfuscation remains poorly understood. Our preliminary study shows that LLM performance significantly degrades on obfuscated code, suggesting a reliance on superficial lexical cues rather than deep semantic reasoning. To address this limitation, we propose CodeSteer, a novel attention steering approach that reallocates model attention toward semantically relevant program elements, including backward slices for output prediction and control-flow paths for execution reasoning. Our method integrates lightweight program analysis with inference-time attention steering to guide LLMs toward the core input-to-output dependencies of a program. Experiments across multiple models and datasets demonstrate that CodeSteer significantly improves performance on obfuscated code, often recovering comparable accuracy to the level of unobfuscated programs. We also show CodeSteer's practical utility through a case study on buffer overflow detection, highlighting its potential for malware/vulnerability analysis and reverse engineering of obfuscated code.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑