arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于共轭对偶的广义凸性与光滑性:深度神经网络的优化理论

Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks

Binchuan Qi

arXiv 2608.09523首次发表:更新:

发表机构

College of Electronics and Information Engineering Tongji University; Zhejiang Yuying College of Vocational Technology(同济大学电子与信息工程学院; 浙江育英职业技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过Legendre函数与凸共轭构建DNN训练的统一优化框架,引入广义凸性与光滑性,推导广义梯度类优化器的收敛速率,定量分析训练影响因素,实验验证理论与训练动态的一致性。

AI 中文摘要

采用随机梯度下降(SGD)及其变体的深度神经网络(DNN)训练取得了优异的经验性能,但经典优化理论未能完全解释这一成功。这一局限源于传统分析依赖可微性、凸性或光滑性等假设,而DNN的目标函数常违背这些假设。本文通过Legendre函数与凸共轭,为DNN训练构建了统一的优化框架。具体而言,我们引入了$\boldsymbol{\textit{H}}(\boldsymbol{\textit{\u03c8}})$-凸性和$\boldsymbol{\textit{H}}(\boldsymbol{\textit{\u03a8}})$-光滑性,将凸与非凸、光滑与非光滑目标统一于同一形式体系,并揭示了广义光滑性与凸性之间的自然对偶关系。基于这些广义性质,我们通过凸共轭引入了广义梯度下降(GD)与广义SGD,从理论上证明广义GD的最优学习率恰好为1,并为两种提出的优化器推导了严格的基于梯度能量的收敛速率。我们进一步将DNN训练重新表述为复合优化问题,证明其收敛性依赖于同时降低梯度能量并控制网络雅可比矩阵的诱导范数。为表征网络架构与训练配置的实际影响,我们引入了梯度相关因子与模型容量风险,并定量分析了架构设计、批量大小及模型容量如何影响训练收敛性。在不同网络架构、数据集、优化器与损失函数上开展的大量实验,验证了我们的理论界,并证明了理论预测与经验训练动态之间的精确一致性。

英文摘要

Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success. This limitation arises because conventional analyses rely on assumptions such as differentiability, convexity, or smoothness, which are often violated by DNN objectives. In this paper, we establish a unified optimization framework for DNN training by generalizing classical convexity and smoothness through Legendre functions and convex conjugation. Specifically, we introduce $\mathcal{H}(ψ)$-convexity and $\mathcal{H}(Ψ)$-smoothness, which unify convex and non-convex as well as smooth and non-smooth objectives within a single formalism and reveal a natural duality between generalized smoothness and convexity. Building on these generalized properties, we introduce generalized gradient descent (GD) and generalized SGD through convex conjugation. We theoretically prove that generalized GD admits an optimal learning rate of exactly $1$, and derive rigorous gradient-energy-based convergence rates for both proposed optimizers. We further reformulate DNN training as a composite optimization problem, demonstrating that its convergence relies on jointly reducing the gradient energy and controlling the induced norm of the network Jacobian. To characterize the practical influences of network architectures and training configurations, we introduce the gradient correlation factor and model capacity risk, and quantitatively analyze how architectural designs, batch size, and model capacity shape training convergence. Extensive experiments across diverse network architectures, datasets, optimizers, and loss functions validate our theoretical bounds and demonstrate precise alignment between our theoretical predictions and empirical training dynamics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑