发表机构
Systrion GmbH; Wilhelm Büchner Hochschule(Systrion 有限公司; 威廉·毕希纳应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型逆向工程能力评估难的问题,提出Reforge方法,通过特定流程构建函数级真实情况并处理对齐不确定性,经实验验证该方法能揭示优化对模型性能的影响,推动不确定性感知基准测试实践。
AI 中文摘要
大语言模型(LLMs)越来越多地应用于逆向工程任务,然而对其能力的评估却缺乏有效手段。现有基准测试将函数级真实情况构建视为已解决的预处理步骤,未披露可靠评估的函数数量。本文提出Reforge,它通过编译、DWARF和句法提取、对齐及反编译从C源构建函数级真实情况,并将对齐不确定性作为八层置信漏斗和三层分层来操作。在受控微基准测试中,跨优化级别高置信度产出下降,未配对比较会因生存偏差高估优化导致的性能衰减。对七个当代LLMs进行的函数命名概念验证评估展示了基础并推动了不确定性感知基准测试实践。
英文摘要
Large language models (LLMs) are increasingly applied to reverse-engineering tasks, and recent threat-intelligence reporting shows them operating inside live offensive-security workflows. Claims about their capability, however, outpace our ability to measure it. Existing benchmarks for LLM-assisted binary analysis treat the construction of function-level ground truth as a solved pre-processing step and report accuracy without disclosing how many functions were reliably evaluable. We argue that the principal obstacle to fair evaluation is not model capability but the reliability of binary-to-source alignment under compiler optimization. This paper presents Reforge, a provenance-tracked pipeline that constructs function-level ground truth from C source through compilation, DWARF and syntactic extraction, alignment, and decompilation, and that operationalizes alignment uncertainty as an eight-gate confidence funnel with three-tier stratification. On a controlled micro-benchmark, high-confidence yield falls from 87.2% to 65.9% across optimization levels, and unpaired comparisons overstate optimization-induced performance decay through survivorship bias. A proof-of-concept evaluation of seven contemporary LLMs on function naming demonstrates the validity of the concept and generally motivates an uncertainty-aware benchmarking practice.
Comments10 pages, 5 figures; accepted for publication to the 23rd International Conference on Applied Computing 2026, Lisbon October 24-26,2026