arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

约束在线凸优化中与数据相关的遗憾值和波利亚克校正

Data-Dependent Regret and Polyak Corrections for Constrained Online Convex Optimization

Wentao Zhang

arXiv 2607.25480首次发表:更新:

AI 中文总结

研究约束在线凸优化问题,通过保留标准论证中省略的量进行更严格分析,提出AdaOGD - PFS自适应步长方法,在保持每轮可行性时实现\(O(\sqrt{G_T})\)遗憾值,实验使遗憾值界提升38% - 43%。

AI 中文摘要

约束在线凸优化需要在满足每轮凸约束的情况下,针对对抗性凸成本最小化遗憾值,这在安全关键应用中是必需的。一种计算高效的方法将在线梯度下降与波利亚克可行性步骤相结合,每轮使用一次约束评估和一个次梯度。虽然该方法在每轮可行性下实现了\(O(\sqrt{T})\)的遗憾值,但我们通过保留标准最坏情况论证中省略的两个量,得出了更严格的、与数据相关的分析。首先,我们用观察到的累积量\(G_T = \sum_t ||\nabla f_t(x_t)||^2\)替换梯度包络\(G_f^2 T\)。其次,我们确定了一个非负的波利亚克校正\(P_T\),它测量由可行性投影引起的累积平方位移,并以负号进入遗憾值界。由此产生的改进\(\Delta_T = (\eta/2)(G_f^2 T - G_T) + P_T/(2\eta)\)总是非负的。我们还提出了AdaOGD - PFS,一种自适应步长方法,在保持每轮可行性的同时实现了\(O(\sqrt{G_T})\)的遗憾值。在球约束和半空间约束问题上的实验将遗憾值界提高了38%到43%,与数据相关的梯度和波利亚克校正都做出了重大贡献。

英文摘要

Constrained online convex optimization requires minimizing regret against adversarial convex costs while satisfying a convex constraint at every round, as needed in safety-critical applications. A computationally efficient method combines online gradient descent with a Polyak feasibility step, using one constraint evaluation and one subgradient per round. Although this method achieves O(sqrt(T)) regret with per-round feasibility, we derive a tighter, data-dependent analysis by retaining two quantities omitted by the standard worst-case argument. First, we replace the gradient envelope G_f^2 T with the observed accumulation G_T = sum_t ||grad f_t(x_t)||^2. Second, we identify a nonnegative Polyak correction P_T that measures the cumulative squared displacement caused by feasibility projections and enters the regret bound with a negative sign. The resulting improvement, Delta_T = (eta/2)(G_f^2 T - G_T) + P_T/(2 eta), is always nonnegative. We further propose AdaOGD-PFS, an adaptive-step-size method that achieves O(sqrt(G_T)) regret while preserving per-round feasibility. Experiments on ball- and halfspace-constrained problems improve the regret bound by 38 to 43 percent, with both data-dependent gradients and Polyak corrections contributing substantially.

CommentsAccepted for publication in Transactions on Machine Learning Research (TMLR)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑