Transformer中的模式形成
Pattern Formation in Transformers
浏览论文内容
中文总结 AI 辅助
本文利用模式形成理论揭示Transformer的归纳偏置,证明完整架构通过放大行波等新模式施加先验,并可通过位置编码与初始化控制以提升学习效率。
中文摘要 AI 辅助
Transformer架构的归纳偏置是什么?现有关于前向传播如何塑造表征的理论,要么考虑Transformer是否逃离秩坍缩,要么证明自注意力驱动令牌趋向于聚类模式。后一种观点源于一个优雅的动力系统视角,但依赖于简化的架构假设,并且无法解释实践中观察到的丰富结构。这留下了一个重要的开放问题:当完整的Transformer逃离秩坍缩时,它如何结构化令牌表征?利用模式形成理论,我们表明Transformer的动力系统视角可以解释位置编码、多头注意力和输出值几何。我们证明,完整的Transformer架构通过选择性地放大一组丰富的先前未报道的模式(包括行波或旋转波等)来施加归纳先验。我们刻画了每个架构组件在控制哪种模式被放大、哪些模式稳定、竞争或共存中的作用。最后,我们表明这些结构可以充当可控的动力先验,有助于学习。通过选择任务对齐的位置编码和权重初始化,我们在受控序列任务以及CIFAR-10上的ConViT中展示了数据效率的提升和优化的加速。
英文摘要
What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Transformers escape from rank collapse or demonstrates that self-attention drives tokens toward cluster patterns. The latter view arises from an elegant dynamical systems perspective, but relies on simplified architectural assumptions, and does not explain the rich structures observed in practice. This leaves a major open question: when a full Transformer escapes rank collapse, how does it structure token representations? Using pattern-formation theory, we show that the dynamical view of Transformers can account for Positional Encoding, Multi-Head Attention, and Output-Value geometry. We demonstrate that a full Transformer architecture imposes an inductive prior by selectively amplifying a rich set of previously unreported patterns, including traveling or rotating waves among others. We characterize the role of each architectural component in controlling which pattern is amplified, which ones stabilize, compete, or coexist. Finally, we show that these structures can act as a controllable dynamical prior that facilitates learning. By choosing both task-aligned positional encoding and weight initialization, we demonstrate improved data efficiency and accelerated optimization on controlled sequence tasks and with ConViT on CIFAR-10.
发表机构
- LIX, École Polytechnique, IP Paris(巴黎综合理工学院LIX实验室,巴黎综合理工学院,巴黎综合理工学院联盟)
- Centre Borelli, ENS Paris-Saclay, Université Paris-Saclay(巴黎萨克雷高等师范学校博雷利中心,巴黎萨克雷大学)
- Centre d’Analyse et de Mathématique Sociales, EHESS, CNRS(社会分析与数学中心,法国社会科学高等研究院,法国国家科学研究中心)
机构由 AI 辅助整理,请以论文原文为准。