发表机构
Statistical Physics of Computation Laboratory, École Polytechnique Fédérale de Lausanne (EPFL); University of Chinese Academy of Sciences; Information, Learning and Physics Laboratory, École Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院统计物理与计算实验室; 中国科学院大学; 洛桑联邦理工学院信息、学习与物理实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文比较核学习器与特征学习器在单指标任务上的情境学习,用复制方法预测误差并绘制相图,揭示架构选择、数据特性与上下文长度对非线性情境学习的影响。
AI 中文摘要
情境学习(ICL)使预训练模型能够在不更新参数的情况下从演示中推断任务。虽然现有理论大多关注线性目标函数,但本文通过在同一单指标任务族上比较两种单层注意力架构来研究非线性情况。核学习器首先通过固定的非线性特征映射处理输入,然后应用线性注意力;而特征学习器则对原始输入应用注意力,随后进行学习到的非线性读出。我们使用复制方法推导了它们在记忆和泛化误差上的预测,并保留了预训练规模、任务池多样性以及训练和推理上下文长度的影响。所得预测在广泛的机制范围内与数值实验紧密匹配。我们的分析得出了相图,描述了当预训练数据量、任务多样性和上下文长度变化时,每种架构何时具有优势。我们进一步识别了两种学习器在上下文长度缩放上的定性差异。这些结果共同阐明了架构选择如何与数据集相互作用并支配非线性情境学习。
英文摘要
In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tasks. A kernel learner first maps inputs through a fixed nonlinear feature map and then applies linear attention, whereas a feature learner applies attention to the original input, followed by a learned nonlinear readout. We derive predictions for their memorization and generalization errors using the replica method, retaining the effects of pretraining size, task-pool diversity, and training and inference context lengths. The resulting predictions closely match numerical experiments across a broad range of regimes. Our analysis yields phase diagrams that characterize when each architecture is advantageous as the amount of pretraining data, task diversity, and context lengths vary. We further identify qualitatively different context-length scalings for the two learners. Together, these results clarify how architectural choices interact with the dataset and govern nonlinear in-context learning.