发表机构
Vy Labs, Inc.(Vy Labs公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对语言模型记忆的精确删除问题,将删除方式分为减法和重放,在不同参数规模的Gemma 3和Kimi Linear混合模型中验证了两种方式的效果,明确了精确删除与记忆表示的关联。
AI 中文摘要
从持久语言模型记忆中实现精确删除,取决于该记忆如何表示记录。可寻址的影响可通过代数减法移除;而由共享循环状态内的后续写入所改变的影响,则需要从写入前的状态进行重建。我们在两个预训练模型上针对明确省略记录的引用,测试了这种区分。首先,我们将Gemma 3的全局注意力层替换为支持向量记忆。在1B参数规模下进行低秩恢复后,减法与保留键值重新拟合在输出下一个token时的中位数KL散度为5.4×10⁻¹⁵,涉及31个支持token的删除,相对于匹配的微调模型,困惑度仅增加2.0%。在触发、重新学习、采样和LiRA攻击下,掩码重新拟合代理与从未被摄入的基准无法区分。在4B和12B参数规模下,这种验证顺序依然存在,但效用成本分别上升至11.2%和44.3%。其次,在48B参数规模的Kimi Linear混合模型中,加法写入允许固定减法,而对角衰减则对应修正后的减法;而delta规则会使记录12%至49%的贡献依赖于后缀。在确定性MLX实现中,检查点回退并重放可删除长达18842个token上下文中的真实临床记录,其输出logits和所有循环状态与从未被摄入的情况完全一致;重放修正操作可实现精确修订。因此,精确删除是记忆表示的一种属性:对可寻址记录执行减法,对纠缠写入执行重放。
英文摘要
An assistant can stop repeating a fact without removing it from memory. To study this difference, we install a support-vector gate in frozen Gemma 3 and record which stored keys and values belong to each exchange. A deletion request excludes the exchange's rows from the long-range readout and recalculates the gate on what remains. We check this operation against an independent refit, then compare it with running the model again on the conversation without the exchange. This second comparison matters because the exchange may already have influenced surviving memory rows. At 4B, the gated model passed checks for recall and feasible deletion on the same six of eight records admitted by the base model, at a perplexity cost under 2%. Admission fell at the smaller and larger checkpoints with the same configuration. The edited memory agreed closely with the local refit on the registered probes, and the model disclosed fewer deleted answers than when simply instructed to forget. However, an attack evaluated separately for each record could still distinguish edited memory from memory that never stored the record. Excluding an exchange's own rows therefore provides a way to edit and audit conversation memory, while leaving a measurable difference from rebuilding it without that exchange. Additional paired studies found no update-speed advantage for the current FP32 proxy and retained-answer matching below half in every tested condition.
Comments23 pages, 4 figures; code and evidence linked on page 1