arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

被抑制而非被抹除:编辑事实的表征痕迹在无权重知识编辑后依然存在

Suppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing

Priyansh Srivastava, Romit Chatterjee

arXiv 2609.18985首次发表:更新:

发表机构

Sirena Ai(Sirena Ai)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过线性探针证明,知识编辑即使行为成功,原始事实仍可从模型表征中解码,表明编辑是抑制而非抹除,GRACE无权重编辑的结果进一步支持该结论。

AI 中文摘要

知识编辑基准验证了局部正确性,即编辑后的模型在接近编辑的提示上是否产生新事实,但不验证原始事实在模型内还有多少可解码。我们使用线性痕迹探针直接研究残余知识:编辑一个事实后,我们询问原始对象是否仍可从模型的隐藏状态中恢复。在GPT-2-XL上,对50个CounterFact编辑应用三种机制不同的编辑器,成功编辑后原始对象仍可线性解码,且远高于随机水平(ROME的探针准确率为0.96,约束微调为0.86,基于记忆的编辑器GRACE为0.79,随机水平为0.50;所有编辑均达到100%的生成式成功)。GRACE的结果最具信息量:GRACE不改变任何基础模型权重,通过外部记忆覆盖事实,但原始对象仍可从底层网络解码,因此残余痕迹不能归因于不完整的权重更新。我们将此解读为证据,表明编辑即使在行为上成功,也只是在表征空间中抑制而非抹除了原始关联。我们还报告了一个再学习节省工具,该工具在我们的设置中表现不可靠,并讨论了原因;我们将其视为一个负面方法论结果而非证据。代码和数据已发布。

英文摘要

Knowledge-editing benchmarks certify local correctness, whether an edited model produces the new fact on near-edit prompts but not how much of the original fact remains decodable inside the model. We study residual knowledge directly with a linear trace probe: after editing a fact, we ask whether the original object is still recoverable from the model's hidden states. On GPT-2-XL, across three mechanistically distinct editors applied to 50 CounterFact edits, the original object remains linearly decodable well above chance after a successful edit (probe accuracy 0.96 for ROME, 0.86 for constrained fine-tuning, and 0.79 for the memory-based editor GRACE, against a chance level of 0.50; all edits reach 100% generation-based success). The GRACE result is the most informative: GRACE changes zero base-model weights, overriding the fact through an external memory, yet the original object is still decodable from the underlying network, so the residual trace cannot be attributed to an incomplete weight update. We read this as evidence that editing, even when behaviorally successful, suppresses rather than erases the original association in representational space. We also report a relearning-savings instrument that did not behave reliably in our setting and discuss why; we treat it as a negative methodological result rather than evidence. Code and data are released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑