随机特征网络中基于随机矩阵理论的二次弱到强泛化
Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory
浏览论文内容
中文总结 AI 辅助
本文利用随机矩阵理论分析两层随机特征网络中的弱到强泛化,证明学生误差按教师误差平方缩放,实现二次改进,并刻画了不同泛化机制间的转变。
中文摘要 AI 辅助
弱到强泛化是指用弱教师模型产生的标签训练的强学生模型能够比教师模型泛化得更好的现象。本文在两层随机特征网络中研究该现象,其中模型强度由其宽度决定。利用随机矩阵理论的工具,我们推导了最优训练的教师模型和通过梯度流训练的学生模型的总体误差的确定性等价式。对于ReLU激活函数和纯球谐目标,在高斯普适性假设下,我们获得了尖锐的渐近结果,显示出二次改进:学生误差按教师误差的平方缩放。这些结果达到了Medvedev等人(2025)的一般下界。我们还分析了学生在更一般的停止时间和支持多个谐波度数的目标下的行为,刻画了弱到强泛化发生的机制,并识别了二次、非二次和无改进之间的转变。
英文摘要
Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools from random matrix theory, we derive deterministic equivalents for the population errors of an optimally trained teacher and a student trained with gradient flow. For ReLU activation and a pure spherical harmonic target, we obtain sharp asymptotics under a Gaussian universality assumption, showing a quadratic improvement: the student error scales as the square of the teacher error. These results attain the general lower bound of Medvedev at al (2025). We also analyze how the student behaves under more general stopping times and targets supported on multiple harmonic degrees, characterizing the regimes in which weak-to-strong generalization occurs and identifying the transition between quadratic, non-quadratic, and no improvement.
发表机构
- McGill University(麦吉尔大学)
- Mila – Quebec AI Institute(Mila-魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。