线性注意力模型能从上下文非线性教师模型中学到什么?
What can linear attention learn from nonlinear teachers in-context?
查看机构详情
- Harvard University(哈佛大学)
- Williams College(威廉姆斯学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究将线性注意力的渐近学习理论扩展到非线性单索引目标,确立非线性-噪声等价性,明确其局限性并为非线性上下文学习提供可处理的研究起点。
中文摘要 AI 辅助
线性注意力是一种可处理的模型,用于理解Transformer中上下文学习的机制。对于线性回归任务,近期的渐近分析已刻画了其学习与泛化行为。我们将该理论扩展到非线性单索引目标:$y=f(x^\top w)+\boldsymbol{\text{ε}}$。主要结果确立了非线性-噪声等价性:线性注意力仅提取$f$的线性厄米特分量,剩余非线性结构作为有效噪声贡献于泛化误差。这种约简使对应线性理论的结果可迁移到非线性任务。我们阐释了其对有限预训练数据的意义,以及当任务多样性增加时从任务记忆到任务泛化的转变。这些结果明确了约简线性注意力模型的局限性,为研究非线性上下文学习提供了可处理的起点。
英文摘要
Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, $y=f(x^\top w)+\varepsilon $. Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of $f$, while the remaining nonlinear structure contributes to the generalisation error as effective noise. This reduction allows results from the corresponding linear theory to be transferred to nonlinear tasks. We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases. These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.