arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Transformer作为跨任务学习者:共享结构驱动上下文学习中的样本效率

Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning

Zhongjie Shi, Rongjie Lai, Alexander Cloninger, Wenjing Liao

arXiv 2609.29060首次发表:更新:

发表机构

Georgia Institute of Technology; Purdue University; University of California, San Diego(佐治亚理工学院; 普渡大学; 加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过覆盖数刻画任务空间复杂度,提出任务识别与评估程序,并显式构建Softmax注意力Transformer,证明共享跨任务结构能提升上下文学习的样本效率,为首个量化一般非线性任务族跨任务复杂度的工作。

AI 中文摘要

Transformer通过在大规模预训练期间联合学习广泛的任务族,并仅凭简短提示即可适应未见过的任务,从而实现了显著性能。然而,对这一现象的严格数学和统计理解仍然有限。本文旨在研究Transformer如何利用共享的跨任务结构,以及这种结构如何影响上下文学习(ICL)的样本复杂度。具体而言,我们通过规定度量下的覆盖数来刻画任务空间复杂度,从而在无需显式参数表示的情况下量化低维跨任务结构。由此产生的覆盖提供了一组锚函数,我们利用这些锚函数引入一种任务识别与评估程序:上下文观测将未见任务定位到锚函数中,并通过聚合相应锚函数在查询点处的评估来预测响应。在近似方面,我们显式构建了一个具有Softmax注意力的Transformer来近似该程序。在泛化方面,我们推导出一个误差界,该误差界将预训练任务数量和提示长度的影响分离。相对于预训练任务数量的缩放由任务空间和输入域的内在维度决定;一旦任务数量足够多,对提示上下文长度的依赖变为与维度无关。据我们所知,这是首个针对一般非线性任务族量化跨任务复杂度,并显式构建利用其低维结构执行ICL的Transformer的工作。我们的理论为相关任务间的联合预训练如何改善上下文泛化提供了定量解释。

英文摘要

Transformers achieve remarkable performance by jointly learning broad families of tasks during pretraining and adapting to unseen tasks from only a short prompt. Yet a rigorous mathematical and statistical understanding of this phenomenon remains limited. This paper aims to study how Transformers exploit shared cross-task structure and how this structure affects the sample complexity of in-context learning (ICL). Specifically, we characterize task-space complexity through covering numbers under a prescribed metric, thereby quantifying the low-dimensional cross-task structure without requiring an explicit parametric representation. The resulting cover provides a set of anchor functions, which we use to introduce a task-identification-and-evaluation procedure: context observations localize an unseen task among the anchor functions, and the response at a query is predicted by aggregating the corresponding anchor function query evaluations. For approximation, we explicitly construct a Transformer with Softmax attention to approximate this procedure. For generalization, we derive an error bound that separates the effects of the number of pretraining tasks and the prompt length. The scaling with respect to the number of pretraining tasks is governed by the intrinsic dimensions of the task space and input domain; once sufficiently many tasks are available, the dependence on the prompt context length becomes dimension-free. To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL. Our theory provides a quantitative explanation of how joint pretraining across related tasks improves in-context generalization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑