非凸复合优化中马尔可夫采样下的方差缩减条件梯度方法
Variance-Reduced Conditional Gradient Methods under Markovian Sampling for Nonconvex Composite Optimization
浏览论文内容
中文总结 AI 辅助
研究非凸复合优化在马尔可夫采样下的问题,提出MC-ALFCG方法,结合动量条件梯度法等技术,通过嵌套平均、耦合和裁剪控制偏差与二阶矩,简化递归,给出不同噪声情况下的样本复杂度,还进行了数值研究。
中文摘要 AI 辅助
我们研究了在固定遍历马尔可夫链的单个轨迹上梯度样本到达时,在紧致凸集上的随机复合非凸优化问题。现有的单轨迹方差缩减理论涵盖光滑无约束目标;我们使用广义Frank-Wolfe间隙解决无投影复合设置。提出了MC-ALFCG,它结合了动量条件梯度方法、耦合封顶多级蒙特卡罗估计和每次迭代裁剪。最深层嵌套平均使用同一轨迹的连续状态,产生均匀的条件偏差\(O(\tau_{\mathrm{mix}}/T)\),耦合通过迭代位移控制梯度差二阶矩,裁剪强制进行自适应分析所需的路径界。我们在\(\sigma^2\mapsto 2\Lambda G_\sigma^2\)和\(L^2\mapsto 2\Lambda L^2\)下将马尔可夫递归简化为独立采样对应,其中\(\Lambda=O(\tau_{\mathrm{mix}}\log T)\)。对于正中心噪声,调优后的方法实现了期望样本复杂度\(\widetilde{O}((\tau_{\mathrm{mix}}^2G_\sigma+\tau_{\mathrm{mix}}^{5/2}G_\sigma^2)\varepsilon^{-3}+\tau_{\mathrm{mix}}^5\varepsilon^{-2})\)。完全无噪声的特殊情况实现了\(\widetilde{O}(\varepsilon^{-2})\)且无混合时间常数,而一个不依赖混合时间的变体实现了\(\widetilde{O}(\tau_{\mathrm{mix}}^6\varepsilon^{-^{3}}+\tau_{\mathrm{mix}}^3\varepsilon^{-2})\)。所有保证都是在固定转移核下的期望。受控数值研究检验了依赖性敏感性、一个非凸复合实例和裁剪行为。
英文摘要
We study stochastic composite nonconvex optimization over a compact convex set when gradient samples arrive along a single trajectory of a fixed ergodic Markov chain. Existing single-trajectory variance-reduction theory covers smooth unconstrained objectives; we address the projection-free composite setting using the generalized Frank-Wolfe gap. We propose MC-ALFCG, which combines a momentum conditional-gradient method with coupled capped multilevel Monte Carlo estimation and per-iteration clipping. The deepest nested average uses consecutive states from the same trajectory, yielding conditional bias $O(τ_{\mathrm{mix}}/T)$ uniformly over the starting state, while coupling controls the gradient-difference second moment through the iterate displacement. Clipping enforces the pathwise bounds needed by the adaptive analysis. We reduce the Markovian recursion to its independent-sampling counterpart under $σ^2\mapsto 2ΛG_σ^2$ and $L^2\mapsto 2ΛL^2$, where $Λ=O(τ_{\mathrm{mix}}\log T)$. For positive centered noise, the tuned method achieves expected sample complexity $\widetilde{O}((τ_{\mathrm{mix}}^2G_σ+τ_{\mathrm{mix}}^{5/2}G_σ^2)\varepsilon^{-3}+τ_{\mathrm{mix}}^5\varepsilon^{-2})$. The exactly noiseless specialization achieves $\widetilde{O}(\varepsilon^{-2})$ with mixing-time-free constants, while a mixing-time-oblivious variant achieves $\widetilde{O}(τ_{\mathrm{mix}}^6\varepsilon^{-3}+τ_{\mathrm{mix}}^3\varepsilon^{-2})$. All guarantees are in expectation under a fixed transition kernel. Controlled numerical studies examine dependence sensitivity, a nonconvex composite instance, and clipping behavior.
发表机构
- Nanjing University of Information Science and Technology(南京信息工程大学)
机构由 AI 辅助整理,请以论文原文为准。