arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

遗忘并非解决之道:序列记忆痕迹编辑中的路径依赖

Forgetting Is Not a Fix: Path Dependence in Sequential Engram Editing

Ferdinand M. Schessl

arXiv 2607.24805首次发表:更新:

AI 中文总结

研究序列记忆痕迹编辑中路径依赖问题,通过在作者参考实现上测试,发现零样本合成与顺序编辑差异大、切割顺序不可互换、层输入协方差随切割漂移、擦除知识会部分恢复,证伪了交换流形假说。

AI 中文摘要

人工智能记忆痕迹(Kwon等人,2026年)将神经科学的四个记忆痕迹标准形式化为权重空间中的约束逆问题并以封闭形式求解:特定概念的记忆痕迹成为线性对象,可一次性提取并进行算术组合。附录F提出了组合记忆状态假说:编辑后的模型存在于“一个交换流形上,其中A和B的整合无论学习顺序如何都能达到一致的平衡”。证据基础是单编辑和双编辑——从材料角度看是单周期测试,其中疲劳积累在结构上不可见。本文将未来工作推迟到对顺序负载下该假说是否成立的“时间动态”问题上。我们在作者自己的参考实现上进行测试,采用他们报告的最佳编辑强度(豆腐α = 0.6,这一选择有利于线性假说),并进行预注册预测,涵盖三个模型负载(两个供应商,两个架构家族)。有四个发现适用于所有三种情况:(1)零样本合成和顺序重新校准编辑的差异为编辑幅度的61 - 71%;(2)切割顺序不可互换,且效果随概念重叠而变化——在一个负载中,切割两个巴黎地标的顺序决定了一个无关的第三个概念是否留存;(3)幸存者的层输入协方差——该方法自身的充分统计量,可视为应变仪——在每个负载的每个留存概念中,随着每次进一步切割单调漂移;(四)被擦除的知识在后续不相关切割下部分恢复。附录F的交换流形假说因此在序列编辑中被证伪;原论文的单编辑结果不受影响。对于作为合规性的遗忘:今天认证的擦除在下一次编辑后不能认证工件。

英文摘要

AI Engram (Kwon et al., 2026) formalizes the four engram criteria of neuroscience as a constrained inverse problem in weight space and solves it closed-form: concept-specific memory traces become linear objects that can be extracted once and combined arithmetically. Appendix F states the Compositional Memory States Hypothesis: edited models live on "a commutative manifold where the integration of A and B reaches a consistent equilibrium regardless of the learning sequence." The evidence base is single and paired edits -- in materials terms, single-cycle tests, in which fatigue accumulation is structurally invisible. Whether the hypothesis holds under sequential load is exactly the "temporal dynamics" question the paper defers to future work. We run that test on the authors' own reference implementation, at their reported best edit strength (TOFU alpha=0.6, a choice favoring the linearity hypothesis), with pre-registered predictions, across three model charges (two vendors, two architecture families). Four findings replicate across all three: (1) zero-shot composition and sequential re-calibrated editing diverge by 61-71% of the edit magnitude; (2) cut order is not interchangeable, and the effect scales with concept overlap -- in one charge the order of cutting two Paris landmarks decides whether an uninvolved third concept survives; (3) the survivors' layer-input covariances -- the method's own sufficient statistics, read as strain gauges -- drift monotonically with every further cut, in every surviving concept, in every charge; (4) erased knowledge partially returns under subsequent unrelated cuts. Appendix F's commutative-manifold hypothesis is thereby falsified for sequential editing; the single-edit results of the original paper are untouched. For unlearning-as-compliance: erasure certified today does not certify the artifact after its next edit.

Comments8 pages, 3 figures. Companion code and data: https://github.com/FerdinandSchessl/engram-seq-note-companion

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑