AI 中文总结
针对优化与混淆导致的反编译代码结构崩溃和语义幻觉问题,提出多阶段LLM框架ReSource,通过解耦词汇、句法、语义层级恢复,在大规模基准上实现83% Top-5准确率,为安全分析提供可靠基础。
AI 中文摘要
逆向工程对软件安全分析和漏洞检测至关重要,反编译(将二进制代码提升为高级伪代码的过程)是该任务的核心。然而,生产级二进制代码是恶劣环境:激进的编译器优化与对抗性混淆会共同破坏控制结构、模糊变量意图并掩盖高级程序逻辑。因此,现有基于大语言模型(LLM)的反编译工具常出现结构崩溃和语义幻觉。本文提出ReSource,首个专为变换无关源代码恢复设计的多阶段LLM框架。为解决这些交织的扭曲,ReSource将二进制与源代码的差异概念化为词汇、句法、语义三个正交层级,并据此解耦恢复过程:首先,为锚定LLM并防止逻辑漂移,它从精心构建的语义扭曲数据库中检索经验先验;其次,为解决控制流扁平化问题,它集成轻量级预测器以重构源代码级结构骨架;最后,上下文词汇推导阶段会优化标识符以恢复可读性。在包含三个优化等级和四种混淆技术的8万余对反编译-源代码函数的大规模基准上评估,ReSource的Top-5源代码检索准确率达83%,平均相似度分数为0.66。在当前最优基线模型(DeGPT、LLM4Decompile和FidelityGPT)严重过拟合或性能下降的情况下,ReSource仍保持稳健的语义可识别性,为下游安全分析提供了可扩展且可靠的基础。
英文摘要
Reverse engineering is essential for software security analysis and vulnerability detection. Decompilation, the process of lifting binaries to high-level pseudocode, is central to this task. However, production binaries are hostile environments: aggressive compiler optimizations and adversarial obfuscation jointly mangle control structures, obscure variable intents, and disguise high-level program logic. Consequently, existing LLM-based decompilation tools frequently suffer from structural collapse and semantic hallucinations. We present ReSource, the first multi-phase LLM framework designed for transformation-agnostic source recovery. To tackle these intertwined distortions, ReSource conceptualizes the binary-to-source discrepancies into three orthogonal tiers, namely lexical, syntactic, and semantic, and decouples the recovery process accordingly. First, to ground the LLM and prevent logic drift, it retrieves empirical priors from a curated Semantic Distortion Database. Second, to resolve control-flow flattening, it integrates a lightweight predictor to reconstruct the source-level structural skeleton. Finally, a contextual lexical deduction stage refines identifiers to restore human readability. Evaluated on a massive benchmark of over 80,000 decompiled-source function pairs across three optimization levels and four obfuscation techniques, ReSource achieves an 83% Top-5 source retrieval accuracy and an average similarity score of 0.66. By maintaining robust semantic identifiability where state-of-the-art baselines (DeGPT, LLM4Decompile, and FidelityGPT) severely overfit or degrade, ReSource provides a scalable and reliable foundation for downstream security analysis.
CommentsPreprint. 11 pages, 3 figures, 3 tables