发表机构
University of Toronto; Vector Institute; New York University; Flatiron Institute; University of British Columbia; Numurho(多伦多大学; 向量研究所; 纽约大学; 熨斗研究所; 不列颠哥伦比亚大学; Numurho)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究分析多类逻辑回归在高维高斯混合模型上的SGD训练动态,揭示顺序学习现象,并推导出计算最优的缩放定律,将理论从线性回归扩展至多类分类。
AI 中文摘要
我们研究了具有大量类别的高维高斯混合模型上多类逻辑回归的训练动态,并建立了在基于梯度的优化下控制交叉熵风险的精确缩放定律。我们表明,学习按类别顺序进行,从最频繁的类别到最不频繁的类别。当类别先验遵循幂律分布时,风险动态分解为三个阶段:初始平台期,直到第一个类别被学习;幂律衰减期,在此期间发生顺序学习;以及最终收敛期。然后,我们分析了在固定计算预算下模型容量如何与优化相互作用。当通过投影到前导主成分来限制有效维度时,风险分解为容量项(保留维度的幂律)和优化项(训练时间的幂律)。优化这一权衡产生了逻辑回归的计算最优缩放定律,并给出了模型大小和训练时间作为计算函数的明确处方。这些结果将理论缩放定律从线性回归扩展到多类分类,同时与在大规模神经网络中观察到的经验缩放定律相联系。
英文摘要
We study the training dynamics of multiclass logistic regression on high-dimensional Gaussian mixture models with a large number of classes and establish precise scaling laws governing the cross-entropy risk under gradient-based optimization. We show that learning proceeds sequentially across classes, from most to least frequent. When the class priors follow a power law distribution, the risk dynamics decompose into three phases: an initial plateau until the first class is learned, a power-law decay regime during which sequential learning occurs, and a final convergence regime. We then analyze how model capacity interacts with optimization under a fixed compute budget. When the effective dimension is restricted via projection onto leading principal components, the risk decomposes into a capacity term (a power law in the retained dimension) and an optimization term (a power law in training time). Optimizing this tradeoff yields a compute-optimal scaling law for logistic regression, with explicit prescriptions for model size and training time as functions of compute. These results extend theoretical scaling laws from linear regression to multiclass classification, while connecting to empirical scaling laws observed in large-scale neural networks.