遗忘注意力:具有认证选择和精确遗忘的可训练支持向量记忆
What a Deletion Certificate Covers, and Where It Expires: Auditable Removal from a Support-Vector Memory
- Vy Labs, Inc.(Vy实验室公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究提出支持向量注意力(SV - 注意力),通过最大间隔记忆及相关技术实现认证选择和精确遗忘。实验表明其在多种任务中有良好表现,如在真实流数据上的召回率提升,在enwik8数据集上有配对改进,展示了可遗忘检索记忆等能力。
AI中文摘要:
注意力可视为基于上下文的在线学习器,但现有测试时记忆无法证明丢弃一个令牌会使输出不变或直接消除其影响。我们引入支持向量注意力(SV - 注意力),它是一种最大间隔记忆,其权重是具有固定盒参数C的一类支持向量机的支持系数。其活动集划分使保留令牌的权重恰好为零,证明了输出保持驱逐;一个可逆增量求解器删除一个令牌以恢复在相同C下不包含该令牌重新训练产生的状态。在fp64实验中,当最优解唯一时,减量和重新拟合可恢复相同的划分,其决策函数匹配到约10^-9的中位数偏差(在学习键上为10^-13);最坏情况下10^-2的情况仅限于病态重复,并且在每个模式下都低于系数衰减。精确路径在自定义反向传播中重用维护的KKT逆。训练使用单独的稳定批近似,不携带精确删除证书;在一个322万参数模型上达到每秒9125个令牌,比MPS softmax参考慢35.8倍。在匹配预算下,认证选择在真实MIMIC - IV流上的罕见项目召回率达到0.86对0.32,恶化小时数保持在0.80对0.05。我们还展示了手术遗忘、精确编辑、患者记录删除以及基于真实句子嵌入的可遗忘检索记忆。在enwik8上,混合模型在七个种子上的平均比特每字符(BPC)为2.178,而匹配状态滑动窗口Transformer为2.383(配对改进8.6%,p = 0.001);三个种子的TinyStories结果在方向上是积极的但不显著(p = 0.057)。
英文摘要:
We study deletion in a context memory that fits a support-vector boundary around stored keys and uses the resulting nonnegative coefficients to weight their values. An exactly zero coefficient lets us remove a key without changing the current normalized readout. In a three-key construction, however, a later admission makes the discarded key active in a solve over the full history. For a positive-weight key, a decremental update targets a refit on the remaining keys at the original coefficient cap, under its stated assumptions. Earlier standalone comparisons show close typical gate-score agreement, while similar keys with different values can yield larger readout differences. The journal extension tests chained deletions and admissions, checking coefficient feasibility and optimality conditions in both returned states and fresh references. The initial audit completes 66 of 80 trajectories and exposes numerical failures. A separately checked wrapper completes all 5,120 scheduled operations at unchanged acceptance thresholds, with eight initialization rescues and 84 update rescues. Its largest readout difference is 0.5273% of retained-value range. These finite results support explicit acceptance and rescue rules for the evaluated memory edits. They do not extend the current-state removal guarantee to arbitrary future admissions.