基于门控奇异值收缩的选择性知识编辑逆转
Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage
浏览论文内容
中文总结 AI 辅助
针对现有知识编辑逆转方法易误删有益编辑的问题,本文提出基于门控奇异值收缩的选择性知识编辑逆转框架,可精准逆转特定编辑同时保留其余编辑,实验验证了其有效性。
中文摘要 AI 辅助
知识编辑为更新大型语言模型中的事实知识提供了高效方式,但恶意编辑可能引入安全风险,因此需逆转不良编辑效果。现有针对参数修改型编辑的逆转方法主要关注全局移除,可能同时抹去应保留的有益编辑。本文研究编辑知识的选择性逆转,目标是逆转特定编辑的事实同时保留其余编辑的事实。基于每个编辑稀疏编码在被编辑矩阵的主导子空间中的假设,我们提出一种基于谱的逆转框架,用于定位被编辑权重的主导奇异子空间内对编辑敏感的组件。在多种设置下的实验表明,我们的方法在逆转选定编辑的同时保留无关编辑的事实方面具有有效性。这些结果表明,不同编辑稀疏编码在主导奇异组件中,且在编辑数量适中时可分离,这使得选择性谱逆转成为定位编辑特定组件并修复被编辑语言模型的有前景方向。
英文摘要
Knowledge editing provides an efficient way to update factual knowledge in large language models. However, malicious edits may introduce safety risks, making it necessary to reverse undesirable editing effects. Existing reversal methods for parameter-modifying edits mainly focus on global removal, which may also erase beneficial edits that should be preserved. In this paper, we study selective reversal of edited knowledge, where the goal is to reverse targeted edited facts while preserving the remaining edited facts. Based on the hypothesis that each edit is sparsely encoded within the dominant subspace of the edited matrix, we propose a spectral-based reversal framework that locates edit-sensitive components within the dominant singular subspace of edited weights. Experiments across multiple settings demonstrate the effectiveness of our method in reversing selected edits while preserving unrelated edited facts. These results suggest that different edits are sparsely encoded within dominant singular components and can be separable when the number of edits is moderate, making selective spectral reversal a promising direction for locating edit-specific components and repairing edited language models.
发表机构
- Zhongguancun Laboratory(中关村实验室)
机构由 AI 辅助整理,请以论文原文为准。