发表机构
University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Echo利用编译作为可信反馈指导迭代搜索,通过领域模型生成候选并逐步修复失配,在函数级基准和Mirai上分别实现2.43倍和2.75倍/7.4倍的精确匹配提升。
AI 中文摘要
神经反编译器可以从二进制文件中恢复可读且可重编译的源代码,但其预测结果仍难以令人信任。匹配反编译通过搜索其重编译后的汇编代码与目标完全匹配的源代码来解决这一问题,从而提供更强的正确性证据。然而,在未知编译配置下,对于优化后的二进制文件,精确匹配仍然具有挑战性。我们提出了Echo,一个基于可信反向翻译的匹配反编译系统。我们的关键见解是,不仅将编译用于验证,还将其作为可信反馈来指导迭代搜索。Echo首先使用一个领域特定模型生成候选程序和编译配置。它重新编译这些候选,度量汇编级相似性,并合成有前景的代码-配置对。剩余的失配随后通过基于规则的改写、神经细化和基于推理的细化逐步修复。我们在函数级基准和Mirai恶意软件二进制文件上评估了Echo。与最强的基线相比,Echo平均产生2.43倍更多的精确匹配,并实现了与真实源代码的最高结构相似性。在Mirai上,Echo匹配的函数数量分别是GPT-5.6和Codex的2.75倍和7.4倍。
英文摘要
Neural decompilers can recover readable and recompilable source code from binaries, but their predictions remain difficult to trust. Matching decompilation addresses this problem by searching for source code whose recompiled assembly exactly matches the target, providing stronger evidence of correctness. However, exact matching remains challenging for optimized binaries under unknown compilation configurations. We present Echo, a matching decompilation system based on trusted back-translation. Our key insight is to use compilation not only for verification, but also as trusted feedback to guide iterative search. Echo first uses a domain-specific model to generate candidate programs and compilation configurations. It recompiles these candidates, measures assembly-level similarity, and synthesizes promising code-configuration pairs. Remaining mismatches are then progressively repaired using rule-based rewriting, neural refinement, and reasoning-based refinement. We evaluate Echo on function-level benchmarks and the Mirai malware binary. Compared with the strongest baseline, Echo produces 2.43x more exact matches on average and achieves the highest structural similarity to ground-truth source code. On Mirai, Echo matches 2.75x and 7.4x as many functions as GPT-5.6 and Codex, respectively.
Comments19 pages, 8 figures