发表机构
State Key Laboratory of Internet of Things for Smart City, University of Macau; RIKEN Center for Advanced Intelligence Project; The University of Melbourne; King Abdullah University of Science and Technology; Mila – Québec AI Institute; McGill University; Google Research; The University of Tokyo(澳门大学智慧城市物联网国家重点实验室; 理化学研究所先进智能研究中心; 墨尔本大学; 阿卜杜拉国王科技大学; 米拉-魁北克人工智能研究所; 麦吉尔大学; 谷歌研究院; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对开放权重LLM被微调滥用获取禁止能力的风险,提出基于梯度密封的先发式遗忘方法,通过将预激活推入负区域阻断梯度通路,实验证明其优于回顾式基线。
AI 中文摘要
开放权重的大型语言模型(LLM)不仅作为固定产品发布,还作为下游微调的基础。然而,这种开放性带来了法律和伦理风险,因为用户可能滥用微调来灌输非法知识或启用敌对操作。因此,模型提供商需要一种发布前的防御措施来应对此类获取,这催生了先发式遗忘(preemptive unlearning)问题。与回顾式遗忘(retrospective unlearning)不同——后者移除固定模型中已存在的能力——先发式遗忘旨在防止在未见过的攻击数据和未来的微调过程中获取这些能力。尽管其具有实际重要性,这一设置在很大程度上仍未得到探索,带来了独特的挑战,因此成为我们工作的核心焦点。我们首先验证了现有的回顾式方法无法提供足够的发布前保护。即使当前输出中抑制了被禁止的能力,被禁止领域的数据仍可通过内部通路诱导梯度,从而实现后续获取。基于这一发现,我们提出了梯度密封(gradient-sealing)原则,通过将相关预激活推入负区域(ReLU系列激活函数在该区域表现出零或接近零的导数)来阻断这些通路。跨多个LLM家族的实验表明,与回顾式基线相比,我们对下游获取具有更强的抵抗力,验证了梯度密封作为发布前保护的有效机制。
英文摘要
Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against such acquisition, motivating the problem of preemptive unlearning. Unlike retrospective unlearning, which removes capabilities already present in a fixed model, preemptive unlearning seeks to prevent their acquisition under unseen attack data and future fine-tuning procedures. Despite its practical importance, this setting remains largely unexplored, presents distinct challenges, and is therefore the central focus of our work. We first verify that existing retrospective methods provide insufficient pre-release protection. Even when forbidden capabilities are suppressed in current outputs, forbidden-domain data can still induce gradients through internal pathways, enabling later acquisition. Motivated by this finding, we propose a gradient-sealing principle that blocks these pathways by pushing relevant pre-activations into the negative region, where ReLU-family activations exhibit zero or near-zero derivatives. Experiments across multiple LLM families demonstrate our stronger resistance to downstream acquisition than retrospective baselines, validating gradient sealing as an effective mechanism for pre-release protection.