arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

代码大语言模型在受保护代码上的鲁棒性与权衡

Robustness and Trade-offs for Code LLMs on Protected Code

Jin Wen, Yuejun Guo, Yujie Ma, Qiang Hu, Maxime Cordy

arXiv 2609.04220首次发表:更新:

AI 中文总结

本研究对7种代码LLM在受保护代码任务开展基于执行的对比实验,发现直接推理混淆代码表现优异,无需显式反混淆,模型能力是核心影响因素,支持基于执行指标评估受保护代码工作流。

AI 中文摘要

代码大语言模型(LLMs)正越来越多地用于可能被故意混淆以实现知识产权保护、抗逆向工程或受控访问的软件制品。在逆向工程和安全分析中,反混淆通常被视为下游分析或推理前的预处理步骤,但其在代码LLM流程中的实用性尚未在不同模型和保护方法间得到系统验证。本文对7种代码LLM在受保护代码翻译与补全任务中开展了基于执行的研究,涵盖C++、Go、Java和JavaScript的源程序、5种混淆方法以及3种推理协议: plain(普通)、obfuscated(混淆)和deobfuscated(反混淆)。结果表明,对混淆代码直接进行推理的表现往往与对恢复后代码的推理相当或更优;在该受控基准中,GPT-4.1和Qwen3-Coder-30B等更高能力的模型在混淆翻译输入上的Pass@1保留了约90%,说明显式恢复通常不必要。同模型的恢复确实能修复部分混淆引发的失败,但也会降低许多原本成功的案例,低能力模型的净损失最大。在所有设置中,模型能力是主要影响因素,而源语言和混淆方法则有次要但一致的影响。总体而言,研究结果支持面向模型的流程设计,表明受保护代码工作流应主要基于基于执行的指标而非仅静态相似性进行评估。

英文摘要

Code large language models (LLMs) are increasingly used on software artifacts that may be intentionally obfuscated for intellectual-property protection, reverse-engineering resistance, or controlled access. In reverse engineering and security analysis, deobfuscation is commonly treated as the preprocessing step before downstream analysis or inference, yet its utility for code LLM pipelines has not been systematically validated across models and protection methods. We present an execution-based study of seven code LLMs on protected-code translation and completion, spanning source programs from C++, Go, Java, and JavaScript, five obfuscation methods, and three inference protocols: plain, obfuscated, and deobfuscated. Our results show that direct inference on obfuscated code often matches or exceeds inference on restored code. In this controlled benchmark, higher-capability models such as GPT-4.1 and Qwen3-Coder-30B retain about 90% Pass@1 on obfuscated translation inputs, indicating that explicit restoration is often unnecessary. Same-model restoration does recover some obfuscation-induced failures, but it also degrades many cases that already succeed, with lower-capability models showing the largest net losses. Across settings, model capability is the primary factor, while source language and obfuscation method have secondary but consistent effects. Overall, our findings support model-aware pipeline design and indicate that protected-code workflows should be evaluated primarily with execution-based metrics rather than static similarity alone.

Commentsaccepted by ISSRE 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑