AI 中文总结
本研究针对LLM辅助C++反编译在STL函数上表现不佳的问题,提出STILL结构化语义接口,在StlBench与HumanEval实验中显著提升了反编译代码的可执行性。
AI 中文摘要
LLM辅助的反编译提升了代码可读性与可重新执行性,但在使用标准模板库(STL)的剥离式C++函数上表现仍不佳。编译、优化及符号剥离会移除或模糊容器类型、库调用结构等源代码级语义,而传统反编译工具的输出往往无法恢复这些信息。我们提出STILL,一种结构化语义接口,可从剥离式控制流图预测函数级STL容器语义,并将其渲染为紧凑提示供LLM优化。在StlBench上,STILL可预测常见容器级STL语义,在稳定字符串与向量切片上取得最强跨数据集结果;在剥离式HumanEval反编译任务中,这些提示使DeepSeek-chat优化达到28.4%的可执行性,而无提示优化为17.4%、原始Ghidra反编译为8.9%;提示效用依赖下游主干模型,反编译专用模型需轻量适配才能从同一接口中获益。
英文摘要
LLM-assisted decompilation improves readability and re-executability, but still underperforms on stripped C++ functions that use the Standard Template Library (STL). Compilation, optimization, and symbol stripping remove or obscure source-level semantics such as container types and library-call structure, while traditional decompiler output often fails to recover them. We present STILL, a structured semantic interface that predicts function-level STL container semantics from stripped control-flow graphs and renders them as compact hints for LLM refinement. On StlBench, STILL predicts common container-level STL semantics, with the strongest cross-dataset results for stable string and vector slices. On stripped HumanEval decompilation, these hints enable DeepSeek-chat refinement to reach 28.4% executability, compared with 17.4% for no-hint refinement and 8.9% for raw Ghidra decompilation; hint utility is downstream-backbone-dependent, with decompilation-specialized models requiring lightweight adaptation to benefit from the same interface.