arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15819cs.LGcs.AI

使用具有线性自注意力的变压器对简单线性回归任务进行上下文学习闭式解

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention

Katsuyuki Hagiwara

首次发表
浏览论文内容

中文总结 AI 辅助

研究利用具有线性自注意力的变压器在简单线性回归任务中进行上下文学习闭式解,通过层归一化近似获得解,以区别于基于梯度下降算法的近似解,并给出了在经 l1 正则化训练的变压器中的实验示例。

中文摘要 AI 辅助

上下文学习是变压器的一个显著特性,近来备受关注。在许多上下文学习研究中,表明变压器能实现线性和非线性回归问题的求解器,大多采用梯度下降算法。但尚不清楚这些实现是否真通过训练获得。本文构建了具有线性自注意力的变压器,在简单回归任务中进行上下文学习最小二乘估计。关键是通过层归一化近似获得闭式(解析)解,而非基于梯度下降算法的近似解。还展示了一个实验示例,当目标输出为最小二乘估计时,我们的实现主要用于经 l1 正则化训练的变压器中。

英文摘要

In-context learning is a remarkable property of transformers and has recently received a lot of interest. In many studies of in-context learning, it has been shown that transformers are capable of implementing solver for linear and non-linear regression problems, in which the most of them implement gradient descent algorithm. However, it is still unclear whether those implementations have actually been acquired through training. In this paper, we construct a transformer with linear self-attention, which in-context learns the least squares estimate in a simple regression task. The point here is that the closed form (analytical) solution is approximately obtained by using layer normalization rather than an approximate solution based on gradient descent algorithm. Then, we show an experimental example, in which our implementation is mainly used in the transformer trained with l1 regularization when the target output is the least squares estimate.

发表机构

  • Faculty of Education, Mie University(三重大学教育学部)

机构由 AI 辅助整理,请以论文原文为准。

↑