arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

$\tilde{O}(\sqrt{T})$ 遗憾与多对数约束违反用于 COCO

$\tilde{O}(\sqrt{T})$ Regret and Polylogarithmic Constraint Violation for COCO

Dhruv Sarkar, Abhishek Sinha

arXiv 2610.03983首次发表:更新:

发表机构

MIT; Tata Institute of Fundamental Research(麻省理工学院; 塔塔基础研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对约束在线凸优化,提出结合连续 Hedge 与消除的方法,实现 $O(\sqrt{T\log T})$ 遗憾和 $O(\log^2 T)$ 累积约束违反,显著降低约束违反至多对数级别。

AI 中文摘要

我们研究具有对抗性凸损失和约束的约束在线凸优化($\mathsf{COCO}$)。在每一轮 $t\in[T]$ 中,学习器从 $d$ 维凸决策集 $\mathcal X$ 中选择 $x_t$,之后自适应对手揭示凸成本函数 $f_t$ 和约束函数 $g_t$。因此,学习器产生成本 $f_t(x_t)$ 和约束违反 $\max\{0,g_t(x_t)\}$,并旨在同时最小化整个时间范围内的遗憾和累积约束违反($\mathsf{CCV}$)。现有算法实现 $O(\sqrt{T})$ 遗憾和 $\widetilde O(\sqrt{T})$ $\mathsf{CCV}$。我们证明在线策略可以实现 $O(\sqrt{T\log T})$ 遗憾和 $O(\log^2 T)$ $\mathsf{CCV}$,将 $\mathsf{CCV}$ 从多项式减少到多对数,同时保持接近最优的遗憾。我们的方法结合了连续 Hedge 与在收缩可行集上的消除。关键观察是,每当 Hedge 分布的均值违反约束时,Grünbaum 不等式保证 Hedge 概率质量的恒定分数被消除。我们使用自适应学习率调度和将存活体积与学习率耦合的势函数,将这种概率质量减少转化为 $\mathsf{CCV}$ 上 $O(\log^2 T)$ 的界限。

英文摘要

We study constrained online convex optimization with adversarial convex losses and constraints ($\mathsf{COCO}$). At each round \(t\in[T]\), a learner selects \(x_t\) from a \(d\)-dimensional convex decision set \(\mathcal X\), after which an adaptive adversary reveals a convex cost function \(f_t\) and constraint function \(g_t\). Consequently, the learner incurs cost \(f_t(x_t)\) and constraint violation \(\max\{0,g_t(x_t)\}\), and aims to simultaneously minimize regret and cumulative constraint violation ($\mathsf{CCV}$) over the entire horizon. Existing algorithms achieve \(O(\sqrt{T})\) regret and \(\widetilde O(\sqrt{T})\) $\mathsf{CCV}$. We show that an online policy can achieve \(O(\sqrt{T\log T})\) regret and \(O(\log^2 T)\) $\mathsf{CCV}$, reducing the $\mathsf{CCV}$ from polynomial to polylogarithmic while retaining near-optimal regret. Our approach combines continuous Hedge with elimination on shrinking feasible sets. The key observation is that whenever the mean of the Hedge distribution violates a constraint, Grünbaum's inequality guarantees that a constant fraction of the Hedge probability mass is eliminated. We use an adaptive learning-rate schedule and a potential function coupling the surviving volume with the learning rate to convert this probability-mass reduction into a bound of \(O(\log^2 T)\) on the $\mathsf{CCV}$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑