arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05370cs.CRcs.AI

当LLM反编译器的重编译结果更多、保留信息更少

When LLM Decompilers Recompile More and Preserve Less

Chang Liu, Edward Raff, Kristopher Micinski

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对LLM反编译器重编译性与行为一致性分离的问题,提出Decompile-Diverge预言机,实验发现候选者存在4.9%的行为分歧,且重编译性与行为一致性可分离。

中文摘要 AI 辅助

反编译是从编译后的机器代码中恢复高级源代码,是漏洞检测、恶意软件分析等安全任务的基础。Ghidra和Hex-Rays等传统反编译器会将无法解析的内容显示为可见占位符,通常生成无法编译或执行的伪代码;基于LLM的反编译器则会输出干净、符合习惯的C代码,目前其评估几乎完全依赖重编译性和重执行性,即输出是否能构建并通过其提供的输入输出测试。我们证明这些指标可能奖励错误路径:一个函数可能能重编译并通过所有提供的测试,但在其他合法输入上会出现分歧,且已披露的漏洞可能从重编译代码中消失,没有任何崩溃的可见痕迹,现有测试套件无法捕捉这两种失败。为解决这一差距,我们提出Decompile-Diverge,一种不依赖固定或手工测试的行为比较预言机:对每个函数,它合成驱动程序,从参考代码生成模糊测试语料库,并对相同输入重新运行反编译后的代码以检测函数行为的变化。在既定LLM反编译语料库上的9种配置的8个系统中,通过所有提供测试的候选者在我们的输入语料库上仍与原始代码存在分歧:总体占4.9%,单个系统最高达13%。在300个真实GitHub库函数和287个基于CVE的函数上,重编译性和行为一致性可能相互分离:性能最强的改进型LLM将Ghidra的构建率从75%提升至90%,但其匹配率从74%降至62%;在已披露的漏洞中,多达十分之一的漏洞在其输出中表现出崩溃缺失。源代码级分析将这种分歧追溯到引入的字段、类型、被调用函数和保护措施,这些取代了传统工具留下的可见未知内容。

英文摘要

Decompilation recovers high-level source from compiled machine code and serves as a foundation for security tasks such as vulnerability detection and malware analysis. Traditional decompilers like Ghidra and Hex-Rays expose whatever they cannot resolve as visible placeholders and often emit pseudocode that will not compile or execute; LLM-based decompilers produce clean, idiomatic C and are now judged almost entirely by recompilability and re-executability: whether the output builds and passes its shipped input/output tests. We show that these metrics can reward the wrong path: a function may recompile and pass every shipped test yet diverge on other legitimate inputs, and a disclosed vulnerability may disappear from the recompiled code with no visible trace of the crash. Neither failure is caught by existing suites. To address this gap, we propose Decompile-Diverge, a behavioral comparison oracle not relying on fixed or hand-crafted tests: for each function it synthesizes a driver, grows a fuzzing corpus from the reference, and reruns the decompiled code on the same inputs to detect changes in the function's behavior. Across eight systems in nine configurations on established LLM decompilation corpora, candidates that pass every shipped test still diverge from the original on our input corpus: 4.9% overall, and as many as 13% for a single system. On 300 real GitHub library functions and 287 CVE-grounded functions, recompilability and behavioral agreement can come apart: the strongest refinement LLM lifts Ghidra's build rate from 75% to 90%, while its Matched rate falls from 74% to 62%; on disclosed vulnerabilities, up to one tenth exhibit Crash Absence in its output. Source-level analysis traces this divergence to introduced fields, types, callees, and guards that replace the visible unknowns traditional tools leave behind.

发表机构

  • Syracuse University(雪城大学)
  • CrowdStrike(CrowdStrike公司)

机构由 AI 辅助整理,请以论文原文为准。

↑