arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11897cs.LG

半直接傅里叶增量注意力:具有构造性分块-WY核的相位控制增量记忆

Semidirect Fourier Delta Attention: Phase-Controlled Delta Memory with Constructive Chunk-WY Kernels

Tiantian Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对线性注意力在状态跟踪和长上下文记忆的局限,提出半直接傅里叶增量注意力(SFDA),通过构造性分块-WY分解,实现精确仿射分块转移等,经数值验证和实验表明其能学习循环记忆,优于禁用相位的KDA基线。

中文摘要 AI 辅助

线性注意力用固定循环状态取代了softmax注意力不断增长的KV缓存,但这种压缩限制了确切状态跟踪和长上下文记忆。我们引入了半直接傅里叶增量注意力(SFDA),它是Kimi Delta注意力的相位控制推广,用块旋转傅里叶控制取代了实对角衰减。通过对特定乘积的构造性分块-WY分解得到了精确仿射分块转移等结果。通过数值验证代数并在玩具状态跟踪实验中表明,SFDA能学习循环记忆,而禁用相位的KDA基线接近随机。融合核和大规模语言模型比较留待未来工作。

英文摘要

Linear attention replaces softmax attention's growing KV cache with a fixed recurrent state, but this compression limits exact state tracking and long-context memory. We introduce \emph{Semidirect Fourier Delta Attention} (SFDA), a phase-controlled generalization of Kimi Delta Attention that replaces real diagonal decay with block-rotational Fourier control: \[ S_t=(I-β_t k_tk_t^*)Λ_tS_{t-1}+β_tk_tv_t^*, \qquad Λ_t=\diag(α_t\odot e^{iθ_t}). \] Our main result is a constructive chunk-WY factorization for products \(A_t=Λ_t-u_tr_t^*\), giving \[ A_t\cdots A_1=Γ_t-Y_tM_tW_t^* \] with rank growth bounded inside fixed chunks. This yields an exact affine chunk transfer, formal stability and complexity bounds, and a compact characterization of phase-plus-low-rank memory. We verify the algebra numerically and show in toy state-tracking experiments that SFDA learns cyclic memory where the phase-disabled KDA baseline remains near chance. Fused kernels and large-scale language-model comparisons are left to future work.

发表机构

  • Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑