发表机构
Edmond and Lily Safra Center for Brain Sciences, The Hebrew University of Jerusalem(耶路撒冷希伯来大学埃德蒙和莉莉·萨弗拉脑科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种限制嵌入维度和头大小为2的最小Transformer框架,实现内部表示全可视化,揭示学习几何隐含算法,并在简单数字序列任务上逐步演示其计算过程。
AI 中文摘要
我们提出了一个用于构建和解释最小Transformer模型的框架。通过将Transformer的嵌入维度和头大小限制为2,我们能够对其内部表示进行完整的二维可视化。嵌入、查询/键/值变换、注意力输出、残差流和决策边界都可以直接看到。我们的核心主张是,学习到的几何结构隐含了一种算法;R^2中点和边界的排列可以被解读为逐步执行的过程。我们在一个简单任务上训练了一个Transformer,该任务要求每当序列中出现'+'运算符时,模型必须产生最近观察到的偶数。训练完成后,我们逐步可视化Transformer计算的每一步。我们展示了模型如何嵌入标记及其在序列中的相应位置,通过Q、K和V矩阵对它们进行变换,利用Q和K表示之间的点积形成注意力矩阵,并使用注意力矩阵选择值,将每个输入标记的表示移动到输出层域中能正确预测下一个标记的区域。我们引入了一套可解释性可视化工具,使这一过程的算法解释变得明确。我们的框架提供了一个教学和实验测试平台,以探索Transformer如何利用信息几何来实现下一个标记预测。
英文摘要
We present a framework for building and interpreting minimal transformer models. By constraining a transformer's embedding dimension and head size to 2, we enable full two-dimensional visualization of its internal representations. Embeddings, query/key/value transforms, attention outputs, residual streams, and decision boundaries can all be seen directly. Our central claim is that the learned geometry implies an algorithm; the arrangement of points and boundaries in R^2 can be read as a step-by-step procedure. We train a transformer on a simple task where it must produce the most recently observed even number whenever the '+' operator appears in a sequence of digits. Once trained, we visually walk through every step of the transformer's computation. We show how the model embeds the tokens and their respective positions in the sequence, transforms them via the Q, K, and V matrices, uses the dot product between the Q and K representations to form the attention matrix, and uses the attention matrix to select values that move the representation of each input token to the region of the domain of the output layer that will correctly predict the next token. We introduce a suite of interpretability visualizations that make the algorithmic interpretation of this procedure explicit. Our framework offers a pedagogical and experimental testbed to explore how transformers use informational geometry to implement next-token prediction.
Comments27 pages, 15 figures, 2 tables. Code and training-dynamics animations: https://github.com/Raneem-mahajne/creating_transformer