发表机构
A*STAR Institute of Advanced Intelligence and Computing; National University of Singapore(A*STAR先进智能计算研究所; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM辅助的二进制到源代码恢复进行知识系统化,构建细粒度分类,通过多维度评估45000个样本,分析设计选择影响,测试同/跨语言恢复能力,为该领域提供关键见解。
AI 中文摘要
有效的源代码恢复对恶意软件分析、漏洞评估和遗留维护等安全应用至关重要。大语言模型(LLM)正在重塑该领域,将范式从基于规则的启发式方法转向从汇编或经典反编译器生成的伪C中进行概率性、高保真的源代码语义恢复。然而,尽管进展迅速,该领域因众多方法分散且评估不统一而存在问题,限制了客观比较。此外,现有工作对嵌入式、物联网架构以及C/C++之外的源代码语言的覆盖有限。本研究中,我们提出首个专门聚焦于LLM辅助的二进制到源代码恢复的知识系统化(SoK)。我们提供以设计为中心的LLM辅助源代码恢复方法的细粒度分类,并使用六个评估维度下的七个关键指标进行系统评估。我们在来自四个标准和五个嵌入式架构、五个优化级别以及符号剥离的45000个测试样本上评估最先进的方法。我们使用受控的内部恢复流水线和三个现成模型,消融研究设计选择对恢复性能的影响,包括输入表示、上下文丰富度、模型规模、迭代和审查角色。最后,我们测试涵盖四种成熟和两种遗留语言的同语言与跨语言恢复能力。我们的系统化和全面评估提供了指导该领域未来方向的关键见解。
英文摘要
Effective source recovery is critical to security applications such as malware analysis, vulnerability assessment, and legacy maintenance. Large Language Models (LLMs) are reshaping this field, shifting the paradigm away from rule-based heuristics to probabilistic and high fidelity semantic recovery of source code from assembly or classical decompiler-derived pseudo-C. However, despite rapid progress, the field suffers from fragmentation across numerous approaches as well as their non-unified evaluations, limiting objective comparisons. Further, existing works have limited coverage of embedded, IoT architectures and source languages beyond C/C++. In this work, we present the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery. We provide a granular design-centric taxonomy of LLM-assisted source recovery methods and systematic evaluations using seven key metrics along six evaluation dimensions. We evaluate state-of-the-art methods on 45,000 test samples derived from four standard and five embedded architectures, five optimization levels, and symbol stripping. We ablate the impact of design choices on recovery performance, including input representation, contextual enrichment, model scale, iteration and review roles using controlled in-house recovery pipelines and three off-the-shelf models. Finally, we test same-language and cross-language recovery capability covering four mature and two legacy languages. Our systematization and comprehensive evaluations provide key insights that guide future directions in this field.