发表机构
Tsinghua University; Shanghai Jiao Tong University(清华大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出意见领袖动力学框架,将词元视为单位球面上的粒子,通过显式和隐式两种机制解释稀疏注意力下词元组内收敛而组间分离的现象,并在多个前沿大语言模型上验证了该多组结构预测。
AI 中文摘要
稀疏注意力在降低全局自注意力二次方计算成本的同时保持了强大的经验性能,但其受限交互如何塑造词元表征的演化在理论上仍未被充分探索。我们将词元建模为单位球面上的粒子,引入意见领袖动力学这一框架,该框架识别出两种机制,通过这两种机制,词元组在内部收敛的同时保持不同的极限方向。在显式模型中,固定的代表者诱导出一个势能,将词元吸引至不同的局部极大值。在隐式模型中,不连通的交互组朝向各自独立的共识方向演化。我们将两种模型均表述为逆向Wasserstein梯度流,并在适当条件下建立了指数收敛性。我们进一步将这些理论预测与启发我们框架的前沿稀疏注意力大语言模型中的词元演化联系起来。在四个基准上,Kimi-K3、MiniMax-M3和DeepSeek-V4-Flash在投影词元表征中始终表现出比稠密注意力模型GLM-4.7-Flash更清晰的聚类分离和更高的聚类得分。这些观察结果支持了所预测的多组结构对已训练的前沿大语言模型的相关性,而有限粒子模拟则展示了理论收敛行为。总之,我们的结果将受限的词元交互与不同的组级吸引子联系起来,提供了关于稀疏注意力如何在支持组内对齐的同时保持组间分离的动力学解释。
英文摘要
Sparse attention reduces the quadratic cost of global self-attention while retaining strong empirical performance, but how its restricted interactions shape the evolution of token representations remains theoretically underexplored. Modeling tokens as particles on the unit sphere, we introduce opinion leader dynamics, a framework that identifies two mechanisms through which token groups converge internally while maintaining distinct limiting directions. In the explicit model, fixed representatives induce a potential that attracts tokens toward distinct local maxima. In the implicit model, disconnected interaction groups evolve toward separate consensus directions. We formulate both models as reverse Wasserstein gradient flows and establish exponential convergence under suitable conditions. We further connect these theoretical predictions to token evolution in frontier sparse-attention LLMs that motivate our framework. Across four benchmarks, Kimi-K3, MiniMax-M3, and DeepSeek-V4-Flash consistently exhibit clearer cluster separation and higher clustering scores than the dense-attention model GLM-4.7-Flash in projected token representations. These observations support the relevance of the predicted multiple-group structure to trained frontier LLMs, while finite-particle simulations illustrate the theoretical convergence behavior. Together, our results connect restricted token interactions to distinct group-level attractors, providing a dynamical account of how sparse attention can support alignment within groups while preserving separation between them.
CommentsCode is available at https://github.com/Jingkun-Liu/Opinion-Leader-Dynamics.git