AI 中文总结
本研究提出无需测试套件的CHISEL框架,结合编译器与模糊测试工具反馈,利用Gemma4:31b LLM从Ghidra伪代码迭代恢复源代码,在120个ExeBench函数上实现高可编译性与可执行率,性能优于现有方法。
AI 中文摘要
反编译旨在从二进制文件中恢复出高级、可编译且语义等价的代码。传统反编译器生成的伪代码可读性差且无法编译,而近期基于大语言模型(LLM)的方法虽能生成可读代码,但存在语义错误。LLM辅助的迭代恢复是新兴研究方向,不过现有工作依赖提供的测试套件实现语义恢复。本研究提出CHISEL,这是一个无需测试套件的框架,用于从Ghidra生成的伪代码中迭代恢复源代码。CHISEL利用编译器(静态分析)和覆盖引导模糊测试工具(差分分析)提供的简单却有效的反馈,辅以丰富的可观测数据用于接地差异检测与反馈、跨迭代差异记忆以及最佳候选保留机制。我们在120个针对x86-64架构编译的ExeBench函数上,结合四种优化级别(O0-O3)、带调试信息与不带调试信息的变体,使用开放权重的Gemma4:31b LLM对CHISEL的可编译性、语义恢复能力、反馈预言机的正确性以及迭代开销进行系统评估。启用所有推荐功能的CHISEL,在平均2.1次迭代下,实现了96.1%的平均可重编译率和79.8%的平均可重执行率;值得注意的是,CHISEL恢复了26%的第一代执行错误,同时其反馈预言机仅错误接受9.4%的候选代码,且在性能上显著优于两项近期的LLM辅助反编译工作。
英文摘要
Decompilation aims to recover high-level, compilable, and semantically equivalent code from binaries. Traditional decompilers produce pseudo-C that is difficult to read and does not compile, while the recent LLM-assisted approaches generate readable, but semantically incorrect code. LLM-aided iterative recovery is an emerging branch of research, but prior works rely on supplied test suites for semantic recovery. In this work, we present CHISEL, a test suite-free framework to iteratively recover source code from Ghidra-derived pseudo-C. CHISEL uses simple yet effective feedback from a compiler (static analysis) and a coverage-guided fuzzer (differential analysis), augmented by rich observables for grounded divergence detection and feedback, cross-iteration divergence memory, and best candidate retention. We systematically evaluate CHISEL for compilation and semantic recovery, feedback oracle soundness, and iteration overhead on 120 ExeBench functions compiled for the x86-64 architecture, across four optimizations (O0-O3), in both stripped and unstripped variants, using the open-weight Gemma4:31b LLM. CHISEL, with all recommended features, achieves an average of 96.1% re-compilability and 79.8% re-executability rates at an average of 2.1 iterations. Significantly, CHISEL recovers 26% of first-generation execution errors. At the same time, CHISEL feedback oracle falsely accepts only 9.4% candidates. Lastly, CHISEL performs significantly better than two recent prior work on LLM-assisted decompilation.
CommentsAccepted for publication in ACM CCS SURE 2026