发表机构
Columbia University; Univ Rennes, Ensai, CNRS, CREST–UMR 9194; Pennsylvania State University(哥伦比亚大学; 雷恩大学、ENSAI、法国国家科学研究中心、CREST-UMR 9194; 宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对约束非凸随机优化,提出仅需一次投影的惩罚近端方法,达到$\widetilde O(\epsilon^{-4})$随机次梯度复杂度,返回精确可行点并保证平稳性度量。
AI 中文摘要
约束非凸优化在现代机器学习中的应用日益增多,例如安全的大语言模型对齐/微调、迁移学习和秩约束的持续学习。常用的方法是投影随机梯度下降(Projected SGD),它要求在每次迭代时进行投影。然而,投影到函数约束上的成本可能远高于一次随机一阶更新。我们研究能否将此操作推迟到非凸随机优化的最后阶段,并且只执行一次。对于弱凸、可能非光滑的目标函数和正则凸约束,我们提出了一种惩罚近端方法,该方法使用$\widetilde O(\epsilon^{-4})$次随机次梯度并仅对约束集进行一次投影。返回的点是精确可行的,并且具有较小的Moreau包络平稳性度量;对于光滑目标函数,同一算法控制投影梯度映射。对于光滑非凸约束,我们假设不可行约束斜率存在全局下界,这允许任意、可能不可行的初始化。一种精确惩罚变体达到了相同的随机预言阶数,并返回一个在$O(\epsilon)$范围内精确可行的$\epsilon$-KKT点。所有中间更新仅使用到简单欧几里得球的投影。这些是预言复杂度保证:单次终端投影的成本是分开的,并且不假设一般非凸集的高效投影算法可由正则性推导得出。
英文摘要
Constrained nonconvex optimization has seen increasing application in modern machine learning, such as safe LLM alignment/finetuning, transfer learning and rank-constrained continual learning. The commonly used approach is Projected SGD which requires projection at every iteration. However, projection onto a functional constraint can cost substantially more than a stochastic first-order update. We study whether this operation can be deferred until the end of nonconvex stochastic optimization, and only do it once. For weakly convex, possibly nonsmooth objectives with regular convex constraints, we give a penalized proximal method that uses $\widetilde O(ε^{-4})$ stochastic subgradients and one projection onto the constraint set. The returned point is exactly feasible and has a small Moreau-envelope stationarity measure; for smooth objectives, the same algorithm controls the projected-gradient mapping. For smooth nonconvex constraints, we assume a global lower bound on the infeasible constraint slope which permits arbitrary, possibly infeasible initialization. An exact-penalty variant attains the same stochastic-oracle order and returns an exactly feasible point within $O(ε)$ of an $ε$-KKT point. All intermediate updates use only projections onto a simple Euclidean ball. These are oracle-complexity guarantees: the cost of the single terminal projection is separate, and no efficient projection algorithm for a general nonconvex set is assumed to follow from regularity.