arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

教师几何形状塑造教师-学生网络中的可学习性

Teacher Geometry Shapes Learnability in Teacher-Student Networks

Kai J. Sandbrink, Flavio Martinelli, Alexander van Meegen, Wulfram Gerstner, Johanni Brea

arXiv 2609.09595首次发表:更新:

发表机构

EPFL; University of Oxford; RWTH Aachen(洛桑联邦理工学院; 牛津大学; 亚琛工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析教师几何对可学习性的影响,发现最大化节点不相似性可提高成功率,并提出调整学习率策略以优化教师-学生网络的训练。

AI 中文摘要

教师-学生系统,其中教师神经网络生成训练标签,以便学生神经网络学习实现相同的函数,被广泛用作研究学习的抽象设置。然而,教师的结构常常被忽视,通常假设参数是随机生成且服从正态分布的。这掩盖了不同教师可学习性之间的显著差异。我们将可学习性形式化为收敛到全局最小值的成功率,它是过参数化、学习算法、学生初始化分布和教师几何的函数。我们既确定了一个最大化节点不相似性的容易分布,也确定了一个最小化节点不相似性的困难分布,并表明这两种分布在广泛的设置范围内以及对于不同的激活函数,会引发显著不同的成功率。为了解释这一差距,我们研究了包含两种不同次优局部最小值的小型神经网络的损失景观:位于数据分布边缘的越界(OOB)最小值和位于内部的内部最小值。假设无限数据和快速读出层,我们将小型网络的损失景观解析地简化为二维,表明内部最小值的吸引区域随教师结构的变化而变化。在较大的网络中,最大程度不相似的教师会引发更多的内部最小值,而最小程度不相似的教师会引发更多的OOB最小值。受这些分析的启发,我们表明,差异化地增加读出层的学习率并减少内部偏置的学习率可以提高成功率。这些发现为缩小教师-学生网络研究与实践中出现的更结构化函数之间的差距迈出了重要一步。

英文摘要

Teacher-student systems, in which a teacher neural network generates training labels so that a student neural network can learn to implement the same function, are widely used as an abstract setting to study learning. However, the structure of the teachers is often overlooked by assuming randomly-generated, normally-distributed parameters. This hides substantial variation in how learnable different teachers are. We formalize learnability as the success rate of converging to the global minimum, as a function of overparameterization, learning algorithm, student initialization distribution, and teacher geometry. We both identify an easy distribution that maximizes node dissimilarity and a hard distribution that minimizes it, and show that these two distributions induce markedly different success rates across a large range of settings and for different activation functions. To explain the gap, we study the loss landscape of small neural networks that contain two distinct kinds of suboptimal local minima, out-of-bounds (OOB) minima at the edge of the data distribution and interior minima within. Assuming infinite data and a fast readout layer, we analytically reduce the loss landscape of small networks to two dimensions, showing that the region of attraction of interior minima changes as a function of teacher structure. In larger networks, maximally dissimilar teachers induce more interior minima, while minimally dissimilar teachers induce more OOB minima. Motivated by these analyses, we show that differentially increasing the learning rate of the readout layer and decreasing the learning rate of the inner biases increases success rates. These findings provide an important step in narrowing the gap between the study of teacher-student networks and more structured functions that arise in practice.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑