arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35913cs.SEcs.AI

UNBIND:面向代码大语言模型的推理时定向引导遗忘

UNBIND: UNlearning By INference-time Directional Steering for Code LLMs

  • Shandong University(山东大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhengyang Shan, Jiayun Xin, Yanjun Lin, Xu Qian, Zhiang Liu, Minghui Xu, Yue Zhang, Qin Hu, Kun Li, Xiuzhen Cheng

AI总结:

UNBIND提出一种推理时定向引导的代码遗忘框架,在不改变模型权重的情况下,通过分离目标与保留代码的隐藏状态方向,实现选择性遗忘,显著降低目标代码复现率并保持编程能力。

AI中文摘要:

代码大语言模型从大规模代码语料库中获取编程能力,但也会记忆那些日后需要移除的实现。当出现版权或安全问题时,需要代码遗忘技术来控制其持续复现。然而,目标代码与保留代码共享计算模式,在遗忘特定实现与保留通用编程能力之间形成了张力。我们提出UNBIND,一个代码遗忘框架,它分别考虑哪些隐藏状态对应于目标代码以及如何抑制其复现。通过为这些目标构建独立的引导方向,UNBIND在保持模型权重不变的情况下实现了推理时的选择性遗忘。我们的评估覆盖了两个代码模型和两个语料库上的十四种基线方法。UNBIND在每种设置下均取得了最高的联合遗忘与效用得分。它将目标代码复现率降低了97.3%至99.1%(以F-BLEU衡量),与原始模型相比,最多少解决两个HumanEval+和六个MBPP+问题。在固定预算下的重复提取测试中,每设置300个目标中,产生至少50个标记精确片段的目标数量从188–262降至0–2。没有提取的片段达到100个标记,平均最佳恢复比率在0.43%至6.45%之间。多语言和相关代码评估进一步表明,在有效遗忘的同时,对有用编程能力的影响有限,支持UNBIND作为选择性代码遗忘的实用方法。

英文摘要:

Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patterns, creating a tension between forgetting specific implementations and preserving general programming ability. We propose \textbf{UNBIND}, a code unlearning framework that separately considers which hidden states correspond to the target code and how to suppress its reproduction. By constructing separate directions for these objectives, UNBIND achieves selective unlearning at inference time while keeping model weights fixed. Our evaluation covers fourteen baselines across two code models and two corpora. UNBIND achieves the highest joint forgetting and utility score in every setting. It reduces target code reproduction by 97.3\% to 99.1\% as measured by F-BLEU, with at most two fewer HumanEval+ and six fewer MBPP+ problems solved than the original models. In repeated extraction tests under a fixed budget, the number of targets yielding exact spans of at least 50 tokens falls from 188--262 to 0--2 out of 300 per setting. No extracted span reaches 100 tokens, and the mean best recovery ratio ranges from 0.43\% to 6.45\%. Multilingual and related-code evaluations further show effective forgetting with limited impact on useful programming capabilities, supporting UNBIND as a practical approach to selective code unlearning.

↑