Voltic:在循环记忆中区分波动性与随机性
Voltic: Distinguishing Volatility from Stochasticity in Recurrent Memory
浏览论文内容
中文总结 AI 辅助
Voltic通过各向异性协方差和输入依赖噪声区分波动性与随机性,在受控回忆和语言模型推理任务中超越基线,使记忆更适应变化。
中文摘要 AI 辅助
循环序列模型必须在每个词元处决定以多强的程度覆盖其记忆。从贝叶斯滤波的角度看,这种写入是卡尔曼更新的增益,由两个方向相反的来源所产生的不确定性决定:波动性,即底层关联变化的速度,以及随机性,即每次观察这些关联的噪声程度。首先,我们展示了门控增量规则记忆的更新是各向同性不确定性下该滤波器的形式。接下来,我们引入Voltic,一种循环记忆,它保持协方差为各向异性,并使两个噪声方差都依赖于输入,因此写入是向量值的,并携带序列中累积的不确定性。密集协方差必须逐词元传播,排除了这些模型所依赖的并行训练。因此,我们给出了两种假设密度近似,对角和准对角,两者都使记忆更新保持增量规则形式,并重用其分块内核。在关联变化且观察被破坏的受控回忆任务中,Voltic领先所有基线。在结合波动性和随机性的任务中,其在两种外推规模上对最强基线的优势都大于在训练规模上的优势。在45M参数的语言模型中,它领先八项推理平均值,并在超出训练上下文长度的情况下实现了比门控基线更高的检索准确率,且吞吐量接近这些基线。因此,从不确定性递归中推导写入使记忆对变化更加敏感。
英文摘要
Recurrent sequence models must decide how strongly to overwrite their memory at each token. Read as Bayesian filtering, this write is the gain of a Kalman update, set by uncertainty from two sources that pull it in opposite directions: volatility, how quickly the underlying associations change, and stochasticity, how noisy each observation of them is. First, we show that the update of gated delta-rule memories is the form this filter takes under isotropic uncertainty. Next, we introduce Voltic, a recurrent memory that keeps the covariance anisotropic and makes both noise variances input-dependent, so the write is vector-valued and carries uncertainty accumulated over the sequence. A dense covariance would have to be propagated token by token, ruling out the parallel training these models depend on. We therefore give two assumed-density approximations, diagonal and quasi-diagonal, both of which leave the memory update in delta-rule form and reuse its chunked kernels. On controlled recall tasks in which associations change and observations are corrupted, Voltic leads all baselines. On the task combining volatility and stochasticity, its margin over the strongest baseline is larger at both extrapolation sizes than at the training sizes. In 45M-parameter language models it leads an eight-task reasoning average and achieves higher retrieval accuracy beyond the training context length than gated baselines, at throughput close to those baselines. Deriving the write from an uncertainty recursion therefore makes memory more responsive to change.
发表机构
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。