发表机构
Harvard University; Bowdoin College; Massachusetts Institute of Technology(哈佛大学; 鲍登学院; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究线性递归特征机在带噪多输出回归中的动力学,证明其学习到的特征矩阵以$O(\sqrt{d/n})$误差逼近理想值,并在真实数据上验证了所学特征。
AI 中文摘要
递归特征机(RFM)通过交替执行两个步骤来学习数据的表示:将预测器拟合到数据集,以及使用平均梯度外积(AGOP)更新该预测器的特征。AGOP与神经网络中特征学习之间的联系,促使线性RFM成为分析训练过程中表示如何演化的简单设置。在此,我们研究了线性RFM在带噪多输出回归中的动力学与统计特性,其中输入数据为各向同性次高斯分布,目标由维度为$d$的低秩教师矩阵生成。我们将线性RFM与迭代重加权最小二乘之间已知的联系,从插值设置扩展到带噪声的岭正则化多输出回归。我们证明,学习到的特征矩阵在每次迭代中始终接近其无限数据理想对应物。即,对于$n$个样本,我们证明特征矩阵中的误差以高概率按$O(\sqrt{d/n})$衰减。在真实世界的文本和单细胞基因表达数据上的实验,展示了这一简单线性模型所学到的特征。
英文摘要
Recursive feature machines (RFMs) learn representations of data by alternating between fitting a predictor to a dataset and updating features of that predictor using the average gradient outer product (AGOP). Connections between AGOPs and feature learning in neural networks motivate linear RFMs as a simple setting for analyzing how representations evolve during training. Here, we study the dynamics and statistics of linear RFM in noisy multi-output regression with isotropic sub-Gaussian input data and targets generated by a low-rank teacher matrix of dimension $d$. We extend the known connection between linear RFM and iteratively reweighted least squares from the interpolating setting to ridge-regularized multi-output regression with noise. We show that the learned feature matrix remains close to its infinite-data ideal counterpart at every iteration. Namely, for $n$ samples, we show the error in the feature matrix decays as $O(\sqrt{d/n})$ with high probability. Experiments on real-world text and single-cell gene-expression data illustrate the features learned by this simple linear model.
Comments51 pages, 8 figures