注意力动力学理论:面向分布外上下文学习
Theory on Attention Dynamics for Out-of-Distribution In-Context Learning
查看机构详情
- Ohio State University(俄亥俄州立大学)
- University of Houston(休斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过分析α型和β型注意力权重的动力学,揭示了Transformer在分布外上下文学习中的误差机制,并探讨了微调对源域性能的影响,为提升OOD泛化提供了理论依据。
中文摘要 AI 辅助
Transformer模型已展现出卓越的上下文学习(ICL)能力,使其无需额外微调即可执行新任务。然而,当遇到偏离训练分布的分布外(OOD)输入时,其性能往往下降,而背后的理论仍缺乏深入理解。为填补这一空白,我们通过所谓的α型和β型注意力权重的动力学相互作用来刻画输入分布偏移下的OOD误差,这两种权重分别代表Transformer在识别正确和错误特征上的置信度。我们的结果表明,每个特征的OOD误差取决于训练特征与OOD特征之间的所有成对交互,并且在某些情况下,Transformer的表现不优于随机猜测。为提升OOD泛化性能,我们进一步研究使用OOD数据进行模型微调的影响,并特别刻画了模型在源域上的遗忘表现。有趣的是,微调后源域性能并非总是下降,这高度依赖于特征偏移的性质:在OOD域上微调持续增强对原始分布中正确特征的识别置信度,而来自其他错误特征的干扰可能增加或减少。我们在合成数据和真实数据上进行了大量实验,以验证上述理论洞见。
英文摘要
Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. However, their performance often deteriorates when encountering out-of-distribution (OOD) inputs that deviate from the training distribution, and the underlying theory remains poorly understood. To fill this gap, we characterize the OOD error under the input distribution shift through the interplay between the dynamics of the so-called $α$-type and $β$-type attention weights, which represent the transformer's confidence in identifying the correct and incorrect features, respectively. Our results indicate that the OOD error for each feature depends on all pairwise interactions between the training features and OOD features, and under certain cases the transformer performs no better than random guessing. To improve the OOD generalization performance, we next investigate the impact of model finetuning with the OOD data, and particularly, characterize the model forgetting performance on the source domain. Interestingly, the performance on the source domain may not always degrade after finetuning, which highly depends on the nature of the feature shift: finetuning on OOD domain keeps enhancing the confidence of identifying correct features from the original distribution, while the interference from other incorrect features may either increase or decrease. Extensive experiments on both synthetic and real data are conducted to corroborate the theoretical insights.