arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20433cs.CLcs.AI

莫尔:让模型为稳健的跨域知识编辑指引自身方向

Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

Jea Kwon, Jiwon Kim, Dong-kyum Kim, Meeyoung Cha

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对语言模型训练后知识编辑核心能力退化问题,提出莫尔方法,通过从模型自身解码分布采样估计保留协方差,无需外部数据,可作插入组件。实验表明该方法能在多模型上扩展保留能力,提升准确率,凸显分布对齐对无损编辑的关键作用。

中文摘要 AI 辅助

语言模型在训练后就固定不变,而世界不断发展。知识编辑成为避免完全重新训练的关键替代方法,但因核心能力退化受限,数学和编程推理能力下降,百科知识记忆仍完好。我们发现这是分布不匹配导致的。基于协方差的编辑器只能保留参考语料库的子空间,无法捕捉训练后如SFT和DPO形成的实际分布。静态外部语料库无法恢复这种变化。我们提出莫尔方法,通过从模型自身解码分布采样直接估计保留协方差C。单个随机词汇令牌种子生成绕过主导采样输出的指令跟随模板,揭示模型内化的更广泛子空间。莫尔无需外部数据,可作为基于协方差编辑器的插入组件。在OLMo-2、Llama-3.1和Qwen-3(7-8B)上,在MEMIT和AlphaEdit下,批量和顺序模式中,莫尔在最易受影响领域持续扩展保留能力,如Qwen3-8B经20000次AlphaEdit批量编辑后,与维基百科基线相比,保留79.9%的GSM8K准确率,而基线仅为10.9%。这些结果表明使保留分布与模型实际分布对齐是无损编辑的关键因素,且模型本身可能是部署系统中该分布最易获取的来源。

英文摘要

While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and programmatic reasoning collapse while encyclopedic recall remains intact. We trace this asymmetric degradation to a distributional mismatch. Covariance-based editors preserve only the subspaces spanned by their reference corpus, but fail to capture the operative distribution shaped by post-training such as SFT and DPO. Static external corpora, including Wikipedia and even the original pretraining mixture, cannot recover this shifted manifold. We propose Moir, which estimates the preservation covariance $C$ directly from the model itself by sampling from its own decoding distribution. Seeding generation with a single random vocabulary token bypasses the instruction-following templates that otherwise dominate sampled outputs, exposing the broader subspaces the model has internalized. Moir requires no external data and serves as a drop-in component for any covariance-based editor, a practical advantage given that the pre- and post-training corpora of most modern LLMs are not publicly accessible. Across OLMo-2, Llama-3.1, and Qwen-3 (7-8B), under both MEMIT and AlphaEdit and in batch and sequential regimes, Moir consistently extends preservation in the most vulnerable domains, most strikingly on Qwen3-8B after 20,000 AlphaEdit batch edits, it retains 79.9% GSM8K accuracy compared to 10.9% with the Wikipedia baseline. These results suggest that aligning the preservation distribution with the model's operative distribution is a key factor in non-destructive editing, and that the model itself may be the most accessible source of that distribution for deployed systems.

发表机构

  • Max Planck Institute for Security and Privacy(马克斯·普朗克安全与隐私研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑