AI 中文总结
本研究揭示顺序知识编辑虽不损害模型整体能力,却显著削弱其区分检索证据可信度的能力,且该效应广泛存在,现有评估标准无法察觉。
AI 中文摘要
知识编辑的评估通常关注编辑后事实是否改变、改写句是否遵循以及无关答案是否保持不变。一个模型可能通过这三项测试,但仍失去一项它们均未衡量的能力:在从未编辑过的事实上,决定应相信哪些检索到的文档。我们通过比较模型对其记忆答案所赋予的log几率与注入段落所主张答案的log几率,在编辑前后,保持查询、段落和两个候选字符串不变,来评分这一仲裁能力。我们最干净的实验组是保守调优的LoRA:在Qwen2.5-7B-Instruct上进行1,000次顺序编辑后,MMLU分数保持不变(精确到小数点后四位),但未触及事实的仲裁量分布范围却下降了36%。选择性预测也随之退化。风险覆盖曲线下的面积增加了0.107,而相同MMLU下范数匹配扰动的增加仅为0.005;模型在最自信的四分之一仲裁决策上的错误率从0.217上升到0.342。这不是能力损失。在五种严重程度上进行随机扰动扫描,当损害严重到将MMLU从0.6275降至0.3725时,其造成的损害(0.088)仍小于MEMIT在0.6050时造成的损害(0.102)。该效应在三个随机种子、两个模型家族、两个数据集、两种探针不相交标准、三种提示模板及改写查询中均保持一致。对保存的权重差进行层消融显示,该效应是分布式的:没有单一层能重现它,而移除任何一层可恢复约一半。在冻结检索器的检索设置下,准确率从0.592降至0.46。一个次要发现可能在实践中更为重要。在我们运行的五个模型与方法配对中,有三个在发布超参数下于1,000次顺序编辑后MMLU降至随机水平,而编辑成功率保持1.00且局部性指标显示正常。从不测量能力的顺序编辑评估无法发现这一点。
英文摘要
Knowledge editing is evaluated on whether the edited fact changed, whether paraphrases follow, and whether unrelated answers stayed put. A model can pass all three and still lose something none of them measures: the ability to decide, on facts that were never edited, which retrieved documents to believe. We score the log odds a model assigns to its remembered answer against the answer an injected passage asserts, before and after editing, holding the query, the passage and both candidate strings fixed. Our cleanest arm is a conservatively tuned LoRA: after 1,000 sequential edits on Qwen2.5-7B-Instruct it leaves MMLU unchanged to four decimal places, yet the spread of the arbitration quantity across untouched facts falls by 36%. Selective prediction degrades with it. Area under the risk-coverage curve rises by 0.107, against 0.005 for a norm-matched perturbation at the same MMLU, and error on the model's most confident quarter of arbitration decisions goes from 0.217 to 0.342. This is not capability loss. Sweeping random perturbation over five severities, damage bad enough to cut MMLU from 0.6275 to 0.3725 produces less harm (0.088) than MEMIT does at 0.6050 (0.102). The effect holds across three seeds, two model families, two datasets, two probe-disjointness criteria, three prompt templates and paraphrased queries. Layer ablation on saved weight deltas shows it is distributed: no single layer reproduces it, and removing any one recovers about half. Under retrieval with a frozen retriever, accuracy falls from 0.592 to 0.46. A secondary finding may matter more in practice. Three of five model and method pairings we ran collapse to chance MMLU at 1,000 sequential edits under published hyperparameters, while edit success stays at 1.00 and locality reads clean. Sequential-editing evaluations that never measure capability cannot see this.
Comments10 pages, 6 figures, 2 tables. Code and experimental artifacts available on request