arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机多梯度下降的改进收敛速率:借助人工智能发现的一个证明

Improved Convergence Rates for Stochastic Multi-Gradient Descent which Close the Gap: A Proof by AI

Lisha Chen

arXiv 2607.18174首次发表:更新:

AI 中文总结

研究光滑非凸随机多目标问题,核心方法是利用随机多梯度下降,通过借助PS度量的利普希茨连续性,将随机多梯度下降收敛速率从\(\widetilde O(T^{-1/4})\)提升到\(\widetilde O(T^{-1})\) ,证明由人工智能辅助发现。

AI 中文摘要

对于光滑非凸随机多目标问题,随机多梯度下降(SMG)根据随机梯度计算目标函数的近似最速公共下降方向。在无偏、方差有界的随机梯度条件下,本文根据平方帕累托平稳性(PS)度量为SMG建立了新的收敛速率。在恒定步长和线性增长的小批量情况下,经过T次迭代后算法输出处的该度量为\(\widetilde O(T^{-1})\)。这改进了Chen等人(2024)在相同设置下得到的\(\widetilde O(T^{-1/4})\)界。速率改进的关键在于利用由多梯度下降算法(MGDA)方向的范数定义的PS度量的利普希茨连续性,而非Chen等人(2024)使用的MGDA方向的\((1/2)\)-赫尔德连续性。该证明是作者在准备研究生课程作业时发现的:ChatGPT 5.4 Thinking Extended根据作者编写的作业解答提示生成了初始证明策略,作者随后对所得论证进行了验证和整理。附录记录了提示内容并总结了学生提交的内容。

英文摘要

For smooth nonconvex stochastic multi-objective optimization, stochastic multi-gradient descent (SMG) computes an approximate steepest common descent direction of the objectives from stochastic gradients. Under standard assumptions, this note establishes a new convergence rate for SMG in terms of the squared Pareto-stationarity (PS) measure. With a constant stepsize and linearly growing mini-batches, the expected squared empirical PS measure at the algorithm's output is $\tilde{O}(T^{-1})$ after $T$ iterations. This improves on the $\tilde{O}(T^{-1/4})$ bound obtained by Chen et al. (JMLR, 2024) for linearly growing batches, without requiring bounded gradients. Here, $\tilde{O}(\cdot)$ suppresses logarithmic factors. The key to the rate improvement is to exploit the Lipschitz continuity of the PS measure, defined by the norm of the multi-gradient descent algorithm (MGDA) direction, rather than the $(1/2)$-Hölder continuity of the MGDA direction used by Chen et al. (2024). For variants of MoCo and MoCo+ with exact MGDA updates and constant batch sizes, the analysis gives expected squared empirical PS rates of $O(T^{-1/2})$ and $\tilde{O}(T^{-2/3})$, respectively. These bounds match the corresponding single-objective momentum rates up to logarithmic factors. The proof was accidentally discovered while the author was preparing homework for a graduate course: ChatGPT 5.4 Thinking Extended generated the initial proof strategy for the main SMG result in response to an author-written homework-solution prompt; the author then verified and reorganized the resulting argument. The appendices document the prompt, summarize the student submissions with LLM disclosures, and present additional proofs and results on momentum-based stochastic MGDA.

Comments20 pages, preprint, add results and proofs for momentum-based methods, improve the literature discussion

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑