arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17236cs.SE

反编译代码的LLM细化记忆下限

A Memorization Floor for LLM Refinement of Decompiled Code

  • School of Electrical Engineering and Computer Science (SEECS), National University of Sciences and Technology (NUST)(电气与计算机工程学院,国立科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Muhammad Asjad

AI总结:

通过项内对照和排列零模型,测量LLM细化反编译代码时从输入与先验中恢复命名的程度,发现恢复真实但独立于输入消融,且零模型有界,未确认任何注册假设。

AI中文摘要:

我们引入了一个记忆下限:一种项内对照,用于区分LLM对反编译器输出的细化中,哪些信息来自其输入,哪些来自其先验知识。先细化一个函数,然后从标识符已被破坏的输入中再次细化它,并测量哪些信息得以保留。由于该比较是项内的,语料库难度不会产生影响;这花费了二十次API调用。应用于在我们的分析计划确定后编写的函数(因此任何已发布的模型都不可能记住它们),它报告了两件事。恢复是真实的:细化后的输出比基于其自身输出词汇构建的、与处理臂匹配的排列零模型高出+0.072至+0.137。但它并不依赖于我们所消融的输入:破坏输入的数据流会使命名增益改变+0.001(95%置信区间[-0.026, +0.026]),而移除类型前缀或置换名称的改变也不会更大。来自另一供应商的第二台细化器,预先注册并给予字节相同的输入,重现了这一结果——十二个对比、两个模型、十二个零模型。可读性始终保持在最高水平,因此读者得不到任何信号。该零模型是有界的,而非绝对的:低于0.056的贡献是不可见的,且消融保留了操作,因此仅从这些操作中得出的命名仍然是一个竞争性的解释。没有注册的假设得到确认,我们完整报告了背后的五个仪器故障,包括一个对处理臂有偏见的重新组装工具,以及一个我们在注册时未检查其是否适用于我们输入的等价性检查器。

英文摘要:

We introduce a memorization floor: a within-item control separating what LLM refinement of decompiler output recovers from its input from what it recovers from its prior. Refine a function, then refine it again from an input whose identifiers have been destroyed, and measure what survives. Because the comparison is within-item, corpus difficulty cannot contribute; it costs twenty API calls. Applied to functions written after our analysis plan was committed, so no released model could have memorized them, it reports two things. Recovery is real: refined output sits +0.072 to +0.137 above an arm-matched permutation null built from its own output vocabulary. But it does not depend on the input we ablate: destroying the input's dataflow changes the naming gain by +0.001 (95% CI [-0.026, +0.026]), and removing type prefixes or permuting names changes it by no more. A second refiner from another vendor, registered in advance and given byte-identical inputs, reproduces this -- twelve contrasts, two models, twelve nulls. Readability stays at ceiling throughout, so a reader is given no signal. The null is bounded, not absolute: contributions under 0.056 are invisible, and the ablation leaves operations intact, so naming from those alone remains a competing reading. No registered hypothesis was confirmed, and we report the five instrument failures behind that in full, including a reassembly harness biased against the treated arm and an equivalence checker we registered without checking it worked on our inputs.

补充信息

↑