发表机构
Xidian University; Ant Group(西安电子科技大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种利用编译器反馈和官方测试修复反编译C代码的工作流程,通过测试门槛反馈提升修复的可审计性,在104个Coreutils二进制上实现87.5%的通过率。
AI 中文摘要
反编译的C代码通常只有在修复后才能重新编译,但仅重编译并不能确立测试观察到的行为。重新编译的命令行二进制文件仍可能错误地解析选项、打印不同的字节或返回不同的退出状态。我们提出了一种利用编译器反馈和相关官方测试来修复反编译C代码的几步工作流程。编译器和链接器的诊断首先指导构建修复。一旦修复后的C代码重新编译成二进制文件,冒烟检查和相关官方测试便会暴露行为差异,以进行语义修复。在对104个Coreutils 9.5二进制文件(具有可用的反编译器导出和确定性精确输出冒烟比较)进行的初步静态增强评估中,91个二进制文件(87.5%)重新编译并通过了测试门槛;9个未能在修复预算内重新编译,4个重新编译但仍未通过测试门槛。结果表明,测试门槛反馈可以使LLM辅助的反编译C代码修复比仅编译恢复更具可审计性。
英文摘要
Decompiled C often becomes recompilable only after repair, but recompilation alone does not establish test-observed behavior. A recompiled command-line binary can still parse options incorrectly, print different bytes, or return a different exit status. We present a few-step workflow for repairing decompiled C using compiler feedback and related official tests. Compiler and linker diagnostics first guide build repair. Once the repaired C recompiles into a binary, smoke checks and related official tests expose behavioral discrepancies for semantic repair. In a preliminary static-enriched evaluation on 104 Coreutils 9.5 binaries with available decompiler exports and deterministic exact-output smoke comparisons, 91 binaries (87.5%) recompile and pass the test gate; 9 do not recompile within the repair budget, and 4 recompile but still fail the test gate. The result suggests that test-gate feedback can make LLM-assisted repair of decompiled C more auditable than compile-only recovery.
Comments5 pages, 2 figures