arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03921cs.AIcs.NE

Transformer革命 第一部分:通过输出-权重互连实现的动态处理

The Transformer Revolution, Part 1: Dynamic Processing through Output-Weight Interconnections

Marco Giunti, Fabrizia Giulia Garavaglia

AI总结:

本文提出Transformer推理阶段的SIDPP计算机制,通过输出-权重互连构建依赖提示的动态变换,其贡献随提示长度增加,还推测人类语言处理或为类似SIDPP的形式。

AI中文摘要:

本文对推理阶段的Transformer提出了新的解读。针对将大型语言模型视为仅复制训练中学习到的统计规律的“随机鹦鹉”观点,我们认为Transformer会构建并应用依赖于提示的变换,其参数在推理阶段生成。我们将这种计算形式命名为SIDPP:序列级交互式动态并行处理。Transformer被解读为一种通过概念转换概念的系统:词元向量是待转换的概念,由矩阵和向量定义的参数化变换是进行转换的概念,这些变换可以是训练时固定的静态形式,也可以是从输入序列生成的动态形式,在机制上对应着简单神经网络组。Transformer的架构创新在于输出-权重互连,即部分网络的输出决定其他网络的权重,以及常规的输出-输入互连。通过这些互连,系统从提示中构建变换并用于修改词元表示。动态处理的贡献随提示长度增加而增大,可能等于或超过静态处理的贡献,这一现象我们称为强提示敏感性。该解读与可解释性、可预测性、可控性以及更小、更可持续系统的设计相关。最后,由于人类神经系统具备实现SIDPP所需的机制,我们认为某种形式的SIDPP原则上可在大脑皮层中神经实现,因此推测人类语言处理本身可能是一种由与Transformer架构高度相似的功能架构产生的SIDPP形式。

英文摘要:

We reinterpret Transformer inference by developing a functionally equivalent mechanical-structural description of its functional architecture. Parameterized transformations of token representations, or transforming concepts, are identified with simple neural networks organized through output-input and output-weight interconnections. This redescription makes explicit an organizational feature that is not equally salient in the standard matrix description: during inference, the outputs of some networks determine the weights, and hence the transformations, of others. These output-weight interconnections generate prompt-dependent dynamic transformations and give rise to Sequence-level Interactive Dynamic Parallel Processing (SIDPP). We show that the number of dynamic parameters grows linearly with prompt length and may become comparable to, or exceed, the number of static parameters fixed through training, a phenomenon we call strong prompt sensitivity. Philosophically, this shifts the conceptual picture of the Transformer from one centered on the static structure acquired through training to one that also treats the prompt-dependent transformations dynamically constructed during inference as constitutive features of its operation. GPT-4.5's recent Turing test results provide a behavioral illustration of this phenomenon. Finally, we identify biological mechanisms morphologically and functionally correspondent to output-weight interconnections, supporting the in-principle neural realizability of SIDPP and motivating Conjecture T: human neural systems may realize a functional architecture relevantly similar to that of the Transformer.

补充信息

↑