半直接傅里叶增量注意力:具有构造性分块-WY核的相位控制增量记忆
Semidirect Fourier Delta Attention: Phase-Controlled Delta Memory with Constructive Chunk-WY Kernels
浏览论文内容
中文总结 AI 辅助
研究针对线性注意力在状态跟踪和长上下文记忆的局限,提出半直接傅里叶增量注意力(SFDA),通过构造性分块-WY分解,实现精确仿射分块转移等,经数值验证和实验表明其能学习循环记忆,优于禁用相位的KDA基线。
中文摘要 AI 辅助
线性注意力用固定循环状态取代了softmax注意力不断增长的KV缓存,但这种压缩限制了确切状态跟踪和长上下文记忆。我们引入了半直接傅里叶增量注意力(SFDA),它是Kimi Delta注意力的相位控制推广,用块旋转傅里叶控制取代了实对角衰减。通过对特定乘积的构造性分块-WY分解得到了精确仿射分块转移等结果。通过数值验证代数并在玩具状态跟踪实验中表明,SFDA能学习循环记忆,而禁用相位的KDA基线接近随机。融合核和大规模语言模型比较留待未来工作。
英文摘要
Linear attention replaces softmax attention's growing KV cache with a fixed recurrent state, but this compression limits exact state tracking and long-context memory. We introduce \emph{Semidirect Fourier Delta Attention} (SFDA), a phase-controlled generalization of Kimi Delta Attention that replaces real diagonal decay with block-rotational Fourier control: \[ S_t=(I-β_t k_tk_t^*)Λ_tS_{t-1}+β_tk_tv_t^*, \qquad Λ_t=\diag(α_t\odot e^{iθ_t}). \] Our main result is a constructive chunk-WY factorization for products \(A_t=Λ_t-u_tr_t^*\), giving \[ A_t\cdots A_1=Γ_t-Y_tM_tW_t^* \] with rank growth bounded inside fixed chunks. This yields an exact affine chunk transfer, formal stability and complexity bounds, and a compact characterization of phase-plus-low-rank memory. We verify the algebra numerically and show in toy state-tracking experiments that SFDA learns cyclic memory where the phase-disabled KDA baseline remains near chance. Fused kernels and large-scale language-model comparisons are left to future work.
发表机构
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。