arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36230cs.LGmath.DS

前馈层在Transformer动力学中的作用

The Role of Feed-Forward Layers in Transformer Dynamics

Thomas Jacob Maranzatto, Semih Akkoc, Sennur Ulukus

首次发表
浏览论文内容

中文总结 AI 辅助

本文从控制论视角研究Transformer中前馈层的作用,证明前馈网络能将令牌引导至共识状态,并扩展到多簇与多头注意力,通过数值实验验证理论并与真实LLM比较。

中文摘要 AI 辅助

我们从控制论的角度研究Transformer中令牌的动力学行为。我们的模型包含自注意力机制之后存在的前馈层,其中自注意力机制被解释为相互作用的粒子系统,而前馈层则被视为独立的控制。我们的主要理论结果确立了前馈网络能够将令牌引导至任意接近共识的状态,无论键、查询和值矩阵如何。我们的结果很容易扩展到多簇收敛以及多头注意力。我们进行了数值实验以验证我们的结果,并将我们理论中的阈值行为与真实世界的大语言模型进行了比较。

英文摘要

We study the dynamical behavior of tokens in transformers from a control-theoretic perspective. Our model includes the feed-forward layer present after the self-attention mechanism, with the self-attention mechanism interpreted as an interacting particle system and the feed-forward layer as an independent control. Our main theoretical result establishes that the feed-forward network can steer the tokens arbitrarily close to consensus regardless of the key, query, and value matrices. Our result are easily extended to convergence to many clusters and to multi-head attention. We conduct numerical experiments to verify our results, and compare thresholding behavior from our theory to real-world LLMs.

发表机构

  • University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑