从潜在表示到连线:大型语言模型上的外科手术式后期编辑
From Latents to Wires: Surgical Post-Editing on Large Language Models
AI总结:
L2W框架通过雅可比透镜定位和反例引导的因果切断,对大型语言模型实施外科手术式后期编辑,精准移除指定语义目标(如身份声明、内容拒绝)而保留其他能力,并成功应用于文本到图像模型。
AI中文摘要:
给定一个大型语言模型(LLM),持有其权重的人能否指定一个语义目标(例如,模型的身份),定位产生该目标的模型组件,并对其进行编辑,使得该目标不再出现,同时保留其他能力?我们将这种对已训练模型的编辑称为后期编辑。我们提出了L2W(从潜在表示到连线),一个对指定语义目标执行外科手术式后期编辑的框架。在定位方面,L2W使用雅可比透镜(J-lens)归因方法,根据语义目标对组件进行评分。在外科手术式移除方面,由于LLM机制具有冗余性(即,一个语义目标可能由多个组件共同产生),L2W运行反例引导的因果切断(CGCC)算法,直到目标不再出现。CGCC首先累积性地关闭模型组件,将目标的每次存续表达视为一个反例,该反例暴露了接下来需要关闭的组件,然后重新打开其中一些组件以保留能力。在一个植入行为水印的受控实验中,L2W成功移除了水印,并且其定位落在了植入所改变的模型区域上。在三种模型配置下,L2W在所有九次运行中移除了模型元数据(例如,身份)的自我声明,并在全部三次运行中移除了成人内容拒绝,且没有留下任何保留目标的残余。L2W进一步在文本到图像模型上组合了两次后期编辑:一次移除了对请求裸体的拒绝,另一次移除了第一次编辑所暴露的裸体渲染。这些结果支持后期编辑作为后期训练的补充:后期训练安装偏好行为,而后期编辑移除指定的不良行为。
英文摘要:
Given a large language model (LLM), can whoever holds the weights name a semantic target (e.g., the model's identity), locate the model components that produce it, and edit them so that the target no longer appears while other capability is preserved? We call such an edit on a trained model a post-edit. We present L2W (latents to wires), a framework that performs surgical post-edits for named semantic targets. For localization, L2W uses Jacobian lens (J-lens) attribution to score components against the semantic target. For surgical removal, because LLM mechanisms are redundant (i.e., a semantic target may have multiple components producing it), L2W runs Counterexample-Guided Causal Cut (CGCC) until the target no longer appears. CGCC first cumulatively closes model components, treating each surviving expression of the target as a counterexample that exposes the next components to close, and then reopens some of them to preserve capability. In a controlled experiment with an implanted behavioural watermark, L2W removes the watermark, and its localization lands on the model region the implant changed. Across three model configurations, L2W removes model-metadata (e.g., identity) self-claims in all nine runs, and adult-content refusal in all three, with no held-out target residual. L2W further composes two post-edits on a text-to-image model: one removes the refusal of requested nudity, and a second removes the nude rendering the first exposes. The results support post-editing as a complement to post-training: post-training installs preferred behaviours, and post-editing removes named unwanted ones.