arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03858cs.LGcs.AI

检索中心的深度学习在增长型非参数神经网络中的应用

Retrieval-Centric Deep Learning in Growing Nonparametric Neural Networks

  • Google(谷歌)
  • MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
  • Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

Maximilian Schlegel, Rajai Nasser, Seijin Kobayashi, Yanick Schimpf, Oliver Sieberling, Robert Obryk, Kazuki Irie, João Sacramento, Johannes von Oswald

AI总结:

本研究提出检索中心的深度学习(RCDL)范式,通过为每个数据点存储键值表示并在推理时检索重组,取代固定权重矩阵,并开发基于函数梯度的核化注意力学习规则,在图像分类和教师-学生任务上展现高效性能,同时与现有优化器建立联系。

AI中文摘要:

我们研究了一种通用的深度学习层,该层不是将任意大小的训练数据压缩为固定大小的权重矩阵,而是在训练期间为每个数据点存储一对新的键值表示,并在推理时通过注意力机制检索和重组这些表示,从而形成一个增长型神经网络(NN)。虽然Irie等人(arXiv:2202.05798)从经典对偶性角度提出了这一观点,该对偶性将深度NN中通过梯度下降训练的任何线性层表示为对训练数据点的线性注意力(LA),但按照他们的建议,用更强大的注意力函数替换LA并非易事:我们表明,将LA情况下的学习规则天真地应用于高级核函数并不能导致原则性的优化。在此,我们填补了这一空白,并基于径向基函数(RBF)和类softmax核函数,为核化注意力层开发了基于函数梯度的学习规则,从而确立了原则性的“检索中心的深度学习”(RCDL)范式。在实验上,我们在图像分类和合成教师-学生学习任务上展示了RCDL的优异性能和高效学习。此外,我们表明,在NN的对偶形式中用高级LA变体(即MesaNet/DeltaNet)替换LA,与最近提出的针对传统固定大小NN的优化器建立了正式联系,为深度学习优化提供了新的视角。

英文摘要:

We investigate a general-purpose layer for deep learning that, instead of compressing arbitrary-size training data into fixed-size weight matrices, stores a new pair of key-value representations for every data point during training, and retrieves and recombines these representations through an attention mechanism at inference time - resulting in a growing neural net (NN). While Irie et al. (arXiv:2202.05798) have put forward this perspective from the classic duality expressing any linear layer in a deep NN trained by gradient descent as linear attention (LA) over the training data points, replacing LA by more powerful attention functions, as they suggest, turns out to be non-trivial: we show that naively applying learning rules from the LA case to advanced kernels does not lead to principled optimization. Here we fill this gap and develop functional gradient-based learning rules for kernelized attention layers, based on radial basis function (RBF) and softmax-like kernels - establishing the principled "retrieval-centric deep learning" (RCDL) paradigm. Empirically, we demonstrate the promising performance and learning-efficiency of RCDL on image classification and synthetic teacher-student learning tasks. Moreover, we show that replacing LA in the dual form of NNs by advanced LA variants, namely MesaNet/DeltaNet, yields a formal connection to recently proposed optimizers for conventional fixed-size NNs, offering a novel perspective on deep learning optimization.

↑