arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

代码大语言模型的记忆诊断应考虑规模因素

Memorization Diagnostics for Code LLMs Should be Scale-Aware

Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djiré, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawendé F. Bissyandé

arXiv 2608.12771首次发表:更新:

发表机构

University of Luxembourg(卢森堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对代码大语言模型记忆诊断的规模敏感性问题,通过分离表征负载与记忆,发现现有探测技术在大规模模型上失效,提出未来评估需解耦二者的方法。

AI 中文摘要

用于代码的大语言模型在多大程度上依赖记忆而非真正的理解仍存在很大争议。现有文献虽常报告存在广泛的记忆现象,但对稠密架构的探测技术进行评估后发现,这类技术在大规模模型上的效用出现严重失效。采用同义词模糊测试或死代码插入等扰动的传统编码器式探测,即便在已知被污染的基准上,也难以暴露大规模模型的记忆情况;依赖对数概率的解码器式探测则表现出相似的性能下降。这些探测的特定失效模式——尤其是为何这类技术会干扰较小模型却无法影响更大模型——促使我们将表征负载与记忆解耦,而非将二者视为单一现象。通过对数值问题应用可逆数学变换,我们分离出这两个因素并揭示:大规模编码器能成功吸收大量表征负载,同时仍收敛于正确的解族。在实际软件工程中,这种适应不同表面形式的能力才是LLM和智能体应用可用性与可泛化性的关键。某个特定解是否在训练中被见过,已成为一个不那么紧迫的问题,因为尽管记忆会提升被污染基准的分数,但将表征负载分离后,我们对一个功能性答案是否原本是记忆所得的关切就值得商榷。因此,未来的评估必须围绕分离这些现象来构建,而非依赖那些悄悄将二者纠缠的方法。

英文摘要

The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense architectures reveals a severe breakdown in their utility at scale. Traditional encoder-style probes using perturbations such as synonym fuzzing or dead-code insertion struggle to expose memorization in scaled models, even on known-contaminated benchmarks, and decoder-style probes that rely on log probabilities show similar performance degradation. The specific mode of failure for these probes, particularly why such techniques disrupt smaller models but fail to impact larger ones, motivates us to untangle representation load from memorization rather than treating them as a single phenomenon. By applying invertible mathematical transforms to numeric problems, we isolate these two factors and reveal that scaled encoders successfully absorb substantial representation load while still converging on the correct family of solutions. In practical software engineering, this ability to adapt to varying surface forms is what truly matters for usability and generalizability in LLM and agentic applications. Whether a specific solution was seen during training becomes a much less pressing question because although memorization inflates scores on contaminated benchmarks, factoring out representation load makes it debatable how much we should truly care if a functional answer was originally memorized. Future evaluations must therefore be built around separating these phenomena rather than relying on methodologies that quietly entangle them.

Comments26 pages, 6 figures, 6 tables. Under review at EMSE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑