发表机构
Sharif University of Technology(谢里夫理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究弱到强泛化在无正则化两阶段线性回归和随机特征模型中的发生条件,提出主动瓶颈原则:数据量或宽度中较稀缺者控制教师误差过滤,有限资源本身即可正则化。
AI 中文摘要
弱到强泛化(W2SG)发生在学生模型在教师模型预测上训练后表现优于教师模型时。我们研究了在完全收敛、无岭回归的两阶段学习下何时发生这种情况,且没有提前停止、没有显式正则化,也不假设学生模型比教师模型更具表达能力。在两阶段线性回归中,教师模型从$n$个带标签样本中拟合,学生模型仅基于教师模型对$m$个新的无标签输入的预测进行训练。尽管两个阶段共享相同的假设类和训练规则,我们证明学生模型恰好当$m$处于一个明确的中间范围时表现优于教师模型:伪标签太少使学生模型缺乏足够信号,太多则使其继承教师模型的噪声。在幂律协方差下,我们以谱衰减和噪声水平的函数形式推导出该范围的闭式表达式,包括改进区域分裂为两个不相交的$m$区间的场景。随后,我们研究了一个随机特征模型,其中学生模型的特征严格多于教师模型,并识别出两个同样由显式阈值给出的区域:一个区域中改进仅发生在$m$处于有界区间内,另一个区域中改进仅发生在学生宽度$N_S$超过显式阈值时。两个区域均由单一的“主动瓶颈”原则控制:$m$或$N_S$中较稀缺的那个决定教师误差被过滤掉多少,而增加另一个资源只会降低估计噪声。这些结果共同表明,有限数据和有限宽度本身可以正则化两阶段学习器,而无需任何显式机制。
英文摘要
Weak-to-strong generalization (W2SG) occurs when a student trained on a teacher's predictions outperforms that teacher. We study when this happens under fully converged, ridgeless two-stage learning, with no early stopping, no explicit regularization, and no assumption that the student is more expressive than the teacher. In two-stage linear regression, a teacher is fit from $n$ labeled examples and a student is trained solely on the teacher's predictions on $m$ fresh, unlabeled inputs. Although both stages share the same hypothesis class and the same training rule, we show that the student outperforms the teacher exactly when $m$ lies in an explicit intermediate range: too few pseudo-labels leave the student without enough signal, too many let it inherit the teacher's noise. Under power-law covariance, we derive this range in closed form as a function of the spectral decay and noise level, including regimes where the improving region splits into two disjoint intervals of $m$. We then study a random-feature model in which the student has strictly more features than the teacher, and identify two regimes, again given by explicit thresholds: one where improvement occurs only for $m$ in a bounded interval, and one where it occurs only once the student width $N_S$ exceeds an explicit threshold. Both regimes are governed by a single "active-bottleneck" principle: whichever of $m$ or $N_S$ is scarcer controls how much teacher error is filtered out, while increasing the other resource only reduces estimation noise. Together, these results show that finite data and finite width can themselves regularize a two-stage learner, with no explicit mechanism doing so.
Comments76 pages