发表机构
University of California at Los Angeles; Comillas Pontifical University(加州大学洛杉矶分校; 科米利亚斯宗座大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从动力系统角度证明选择性状态空间模型的递归机制与Transformer注意力一样驱动令牌达成共识,并指出输出门调节共识程度。
AI 中文摘要
选择性状态空间模型(SSMs)近来已成为Transformer的一种引人注目的替代方案,将具有竞争力的性能与大幅提升的推理效率相结合。在每个SSM层中,一系列隐藏状态通过递归进行传播,混合不同令牌的信息。尽管使用了不同的机制,这种混合所起的作用类似于Transformer中的注意力。事实上,近期研究表明,这两种架构可能比它们最初看起来更接近,因为这种递归允许一种类似于线性注意力的表述。在Transformer中,已知注意力会驱动令牌聚类,即达成共识,并在极限情况下坍缩到单一方向。因此,我们提出疑问:SSM核心的递归是否像Transformer中的注意力一样驱动令牌达成共识?为回答此问题,我们从动力系统角度审视SSM,将令牌在层间的演化建模为常微分方程。通过利用输入到状态稳定性论证,我们建立了共识平衡点的局部指数稳定性,并刻画了其时变权重矩阵的吸引域,这一设置是先前结果未涉及的。我们由此表明,SSM与Transformer之间的相似性确实更深:SSM核心的递归聚合令牌的方式与注意力相同。在预训练的Mamba-2模型上的数值实验指出,输出门是调节这种共识程度的组件,防止令牌完全达成共识。
英文摘要
Selective state space models (SSMs) have recently emerged as a compelling alternative to transformers, combining competitive performance with substantially improved inference efficiency. At each SSM layer, a sequence of hidden states are propagated by a recurrence, mixing information of different tokens. Despite using a different mechanism, this mixing plays a role analogous to attention in transformers. In fact, recent works have shown that the two architectures may be closer than they first appear, as this recurrence admits a formulation akin to linear attention. In transformers, attention is known to drive the tokens to cluster, i.e., to reach consensus, collapsing in the limit to a single direction. Thus, we ask: does the recurrence at the core of SSMs drive the tokens to consensus, as attention does in transformers? To answer this question, we take a dynamical systems perspective on SSMs, modeling the evolution of tokens across layers as an ordinary differential equation. By exploiting input-to-state stability arguments, we establish local exponential stability of the consensus equilibria and characterize their domain of attraction for time-varying weight matrices, a setting not addressed by previous results. We thereby show that the resemblance between SSMs and transformers does run deeper: the recurrence at the core of SSMs aggregates tokens just as attention does. Numerical experiments on a pretrained Mamba-2 model point to the output gate as the component that regulates the extent of this consensus, preventing the tokens from reaching it in full.