发表机构
Cairo University Computer Engineering Department, Faculty of Engineering(开罗大学计算机工程系,工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对Dart AOT二进制文件神经反编译,在新基准上用三个指标评估六种微调模型变体,发现无微调配置能显著提升pass@k,存在跨语言干扰及指标差异,指出汇编序列长度是任务难度关键预测指标,贡献新基准等并明确pass@k为主要评估指标。
AI 中文摘要
神经反编译作为一个代码生成问题受到越来越多的研究,但其评估方法对于现代语言仍不完善。我们对Dart提前(AOT)神经反编译的微调有效性和指标有效性进行了系统的实证研究。在新的154任务HumanEval-Dart基准上,使用CodeBLEU、compile@k和pass@k这三个指标评估了三种基础架构(4B - 8B参数)的六个微调模型变体。我们的研究得出三个主要发现:一是没有微调配置能在统计上显著提高pass@k;二是Swift训练的跨语言干扰在4B时非常显著,但在8B时与零无统计学差异;三是指标存在差异。错误分析表明汇编序列长度是任务难度的最强预测指标。我们贡献了HumanEval-Dart基准、适用于Dart的CodeBLEU,并证明pass@k必须是神经反编译的主要评估指标。
英文摘要
We present an execution-based evaluation of neural decompilation for Dart ahead-of-time binaries and an audit of what its scores measure. Across six archived adapter-baseline comparisons, paired tests of pass@k at k = 1, 5, and 10, with Holm adjustment over 18 endpoints, identify functional regressions in both Qwen3-8B adapters at every k. The other four comparisons are inconclusive. On 141 reference-certified, contract-valid tasks, three independently trained graph-prefix systems score the same candidates. Best CodeBLEU has modest association with pass@10 ($ρ$ = .218-.246), compile@10 has weak association ($ρ$ = .072-.082), and only 21.0-23.3% of compiling candidates pass. A paired single-seed intervention that removes semantic names and related cues, while retaining types, arity, and instruction content, reduces coverage from 42/154 to 7/154 tasks. Matched graph perturbations show no detectable degradation under the semantic contract (six-test Holm p >= .750); instruction-use attribution remains unresolved. Across five decoding seeds on MF-174, the baseline solves 4.8 tasks on average, 15 at least once, and one in every seed. We recommend certifying references, aligning metrics on shared candidates, separating metadata from binary input, repeating sampling, and preserving provenance. The released capsule supports integrity checks and replay of archived outcomes.
CommentsUnder review at ACM Transactions on Software Engineering and Methodology (TOSEM) after getting a major revision. This is the preprint