arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08922cond-mat.dis-nncond-mat.stat-mechcs.LG

自注意力中的聚类吸引子流形与动力学凝聚

Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention

Qucheng Gao, Zuyi Yang, Xiao Chen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究在最小归一化自注意力动力学中,以重叠间隙为核心量,揭示跨聚类注意力的指数抑制机制,发现注意力锐度超阈值时聚类态成核的动力学注意力凝聚相变。

中文摘要 AI 辅助

Transformer层生成依赖状态的交互网络:token表征决定注意力矩阵,注意力矩阵又更新表征。我们在最小归一化自注意力动力学中研究该反馈,识别重叠间隙为热力学极限下支配其吸引子结构的核心量。当token形成内部对齐的聚类,且它们与同聚类成员的相似度相比与其他所有聚类的相似度超出非零量时,跨聚类注意力随维度增加呈指数级抑制。该机制产生从几个宏观聚类到广泛微观碎片化的高维聚类不动点流形,还控制其对扰动的稳定性。从非结构化高斯态出发,我们发现仅当注意力锐度超过有限阈值时,聚类态才会从弥散背景中成核,引发动力学注意力凝聚相变。

英文摘要

Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations. We study this feedback in a minimal normalized self-attention dynamics and identify the overlap gap as the central quantity governing its attractor structure in the thermodynamic limit. When tokens form internally aligned clusters and their similarity to members of the same cluster exceeds that to every other cluster by a nonvanishing amount, inter-cluster attention is exponentially suppressed as the dimension increases. This mechanism produces a high-dimensional manifold of clustered fixed points, ranging from a few macroscopic clusters to extensive microscopic fragmentation, and also controls their stability against perturbations. Starting from an unstructured Gaussian state, we find that clustered states nucleate from the diffuse background only above a finite threshold in attention sharpness, giving rise to a dynamical attention-condensation transition.

发表机构

  • Boston College(波士顿学院)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑