arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你的Transformer可以同时持有两个想法:LLMs中线性叠加的证据

Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina

arXiv 2609.29845首次发表:更新:

AI 中文总结

本文提出叠加线性假设,证明LLMs在输入线性组合时输出叠加分布,通过轻量级微调恢复线性,并引入引导解码实现单次前向传播生成两个连贯续写。

AI 中文摘要

尽管大型语言模型(LLMs)依赖于高度非线性的组件,但在本工作中,我们证明了它们表现出基本的线性:当来自不同文本流的输入被线性组合时,模型输出的各个下一个词元分布的叠加。我们将其称为“叠加线性假设”。我们提供了证据表明,叠加是Transformer架构的内在属性,而不是训练中涌现的结果;事实上,我们观察到随着预训练的进行,叠加往往会减弱。然而,我们证明了通过轻量级微调可以显著恢复线性,大大减少预测的下一个词元分布与各个下一个词元分布的平均值之间的差异。最后,我们引入了一种引导解码过程,该过程解开叠加的输出,使得单次前向传播能够同时生成两个连贯的续写。

英文摘要

While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the \textit{Superposition Linearity Hypothesis}. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑