发表机构
University of Freiburg; Microsoft Research; University of Tübingen; MPI-IS Tübingen; ELLIS Institute Tübingen; EPFL; Jülich Supercomputing Center (JSC); LAION; PriorLabs(弗莱堡大学; 微软研究院; 图宾根大学; 马克斯·普朗克智能系统研究所图宾根; 图宾根ELLIS研究所; 洛桑联邦理工学院; 于利希超级计算中心; LAION; PriorLabs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 Complex KDA (CKDA),通过扩展 KDA 参数范围实现二维旋转,在保持效率的同时达到 DeltaProduct2 的表达能力,并在状态跟踪和语言建模中表现优异。
AI 中文摘要
基于 delta 规则的线性 RNN 能够实现高效的序列建模,但其带有低秩修正的线性更新限制了其表达能力。先前的工作表明,在单次循环更新中组合两次 delta 规则转换可以模拟二维旋转,但这相比单次转换增加了更新的秩和成本。我们证明 Kimi Delta Attention (KDA) 可以通过将单次 delta 规则变换与其通道门提供的第二次反射相结合来实现二维旋转。这需要结合两种现有的范围扩展来扩展 KDA 的参数范围:允许门控在 $[-1,1]$ 内,delta 规则系数 $\beta$ 在 $[0,2]$ 内。我们将由此产生的模型称为 Complex KDA (CKDA)。它保持了 KDA 的稳定性和效率,其转换保持对角加秩一且非扩张,同时达到了 DeltaProduct$_2$ 的状态跟踪表达能力。我们刻画了 CKDA 的表达能力,并证明了每个正交对角加秩一矩阵恰好是一个 CKDA 转换矩阵。单个 CKDA 层可以跟踪每个与 $\mathrm{SO}(3)$ 的子群同构的有限群,并且许多状态跟踪结果表明,与其他对角加秩一线性 RNN 相比,CKDA 使用的层数少一层。在实验上,结合这两种扩展在 $S_3$、$S_4$ 和周期性音频延续上测试的 KDA 范围设置中产生了最强的长度外推效果。在语言建模中,CKDA 优于 Transformer 和其他线性 RNN,获得了与 KDA 基线相似的结果,并显示出有前景的扩展行为。我们的代码在此 https URL 开源,模型可在此 https URL 获取。
英文摘要
Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that composing two delta-rule transitions in a single recurrent update can model a 2D rotation, but this increases the rank and the cost of the updates compared to a single transition. We show that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate. This requires extending the parameter ranges of KDA by combining two existing range extensions: allowing gates in $[-1,1]$ and the delta-rule coefficient $β$ in $[0,2]$. We call the resulting model Complex KDA (CKDA). It preserves KDA's stability and efficiency, with transitions that remain diagonal-plus-rank-one and non-expansive, while reaching the state-tracking expressivity of DeltaProduct$_2$. We characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix. A single CKDA layer can track every finite group isomorphic to a subgroup of $\mathrm{SO}(3)$, and many state-tracking results use one fewer layer for CKDA compared to other diagonal-plus-rank-one Linear RNNs. Empirically, combining both extensions yields the strongest length extrapolation among tested KDA range settings on $S_3$, $S_4$, and periodic audio continuation. In language modeling, CKDA outperforms Transformers and other linear RNNs, obtains similar results to a KDA baseline, and shows promising scaling behavior. Our code is open source at https://github.com/OpenEuroLLM/ComplexKDA and our models are available at https://huggingface.co/collections/openeurollm/complexkda.