早期梯度下降动力学下逻辑回归的非渐近隐式偏置
Non-asymptotic implicit bias of logistic regression at early-stage gradient descent dynamics
浏览论文内容
中文总结 AI 辅助
本研究揭示早期梯度下降动力学中逻辑回归参数向量与最大间隔方向弱对齐的机制,给出对齐迭代次数的紧上界,为理解训练时长与泛化性能的关联提供理论支撑。
中文摘要 AI 辅助
梯度下降在现代机器学习中备受关注,其重要性远超单纯的优化问题。优化过程中产生的隐式偏置虽未被学习目标编码,却常能避免模型对虚假模式过拟合。线性分类器的最大间隔隐式偏置便是典型例子,该性质已针对指数尾损失函数得到广泛证实:即便给定数据集已线性可分,参数向量仍会沿梯度下降动力学渐近演化至最大间隔方向,这一现象印证了“训练时间越长,泛化性能越好”的常见经验观察。然而,最大间隔收敛是渐近现象,且其渐近收敛速率远慢于纯凸优化。即便如此,在远少于渐近速率的迭代次数内,沿梯度下降动力学演化的参数向量通常会与最大间隔方向呈正相关(虽非完全一致)。本研究从新视角探究这一经典问题,旨在理解这种早期对齐现象的机制。理论结果表明,在$O(\text{exp}(\text{exp}(-\boldsymbol{\u03b4})))$次迭代内,参数向量会与最大间隔方向弱对齐,其中$\boldsymbol{\u03b4}>0$为允许的对齐误差,且该上界被证明是紧的。通过追踪径向和切向流,本研究的证明直接基于数据集几何分析对齐动力学,摒弃了渐近展开,这是实现更快弱对齐的关键见解。
英文摘要
Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization. Implicit bias emerging from optimization, though not being encoded by the learning objective, often prevents from overfitting to spurious patterns. A typical instance is the max-margin implicit bias of a linear classifier, widely established for exponentially tailed loss functions. Even after having a given dataset separated, the parameter vector continues to evolve towards the max-margin direction asymptotically along the gradient descent dynamics. This phenomenon corroborates a frequent empirical observation of "train longer, generalize better." However, the max-margin convergence is an asymptotic phenomenon, and what is worse, this asymptotic convergence rate is significantly slower than pure convex optimization. Even so, the parameter vector along gradient descent dynamics commonly correlates with the max-margin direction positively (though not exactly) within considerably fewer iterations than the asymptotic rate. By shedding another light on this classical problem, this work aims to understand the mechanism of this early-stage alignment phenomenon. Our theoretical results demonstrate that the parameter vector weakly aligns with the max-margin direction within $O(\exp(\exp(-δ)))$ iterations, where $δ>0$ is the permissible alignment error, which is shown to be tight. By tracking the radial and tangential flows, our proof operates on the alignment dynamics directly with dataset geometry and gets rid of the asymptotic expansion, which is a key insight to establishing faster weak alignment.
发表机构
- The Institute of Statistical Mathematics(统计数理研究所)
- Tohoku University(东北大学)
- RIKEN AIP(理化学研究所人工智能项目)
机构由 AI 辅助整理,请以论文原文为准。