arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06921cs.LGcs.AImath.OC

带噪声约束值的约束在线学习

Constrained Online Learning with Noisy Constraint Values

Vaneet Aggarwal

首次发表
浏览论文内容

中文总结 AI 辅助

针对带噪声约束值的约束在线凸优化,提出LEDGER算法,在无Slater条件下实现遗憾与预算违例的权衡,并给出动态遗憾保证。

中文摘要 AI 辅助

我们研究了在约束值和梯度通过无偏噪声观测时,具有对抗性约束的约束在线凸优化问题。标准差为$\sigma$的高斯值噪声,即使在梯度已知的情况下,也会在期望遗憾和期望硬违例的最大值上产生$\Omega(\min\{\sigma,1\}T/\log^7T)$的最坏情况下的下界。这排除了对于固定$\delta>0$和固定正噪声水平,任何联合$O(T^{1-\delta})$的保证。因此,我们研究预算违例:在固定时间范围内任何窗口内的最大累积超支。我们引入了\LEDGER,它在一个非负余额中跟踪观测到的净消耗,并在当前反馈噪声之前设置约束权重。在常见的可行性和条件有限方差反馈下,对于固定问题参数,\LEDGER\\ 实现了$O(\sqrt T/V)$的期望遗憾和$O(\sqrt V\\,T^{3/4}+\sigma\sqrt T)$的期望预算违例,其中$V\in[T^{-1/2},1]$。这给出了在$V=1$时的$(O(\sqrt T),O(T^{3/4}))$和在$V=T^{-1/6}$时的$(O(T^{2/3}),O(T^{2/3}))$,无需Slater条件。预算聚焦端点$V=T^{-1/2}$给出$(O(T),O(\sqrt T))$。相同的更新对于可预测的可行比较器路径,在无需常见可行性或路径长度输入的情况下,产生$O((1+E[P_T])\sqrt T/V)$的期望动态遗憾。其预算界限反而依赖于最短可行路径,最多相差一个维度因子。

英文摘要

We study constrained online convex optimization with adversarial constraints and conditionally unbiased, finite-variance observations of constraint values and gradients. Under common feasibility, our \LEDGER\ algorithm attains $O(\sqrt T)$ expected regret and $O(\sqrt{T\log(eT)})$ expected budget violation, the largest cumulative overspend over any window. It uses a reflected exponential potential, clipped signed observations, and predictable adaptive regularization, with one feedback triple and one projection per round. Neither a Slater condition, independence between feedback channels, nor an absolute constraint-value bound is needed. A Gaussian testing lower bound proves that the budget rate has optimal horizon dependence under square-root regret at fixed positive noise, including the logarithm. The same obstruction holds for terminal violation, so the logarithm is not a cost of maximizing over windows; an $O(\sqrt T)$ budget bound instead forces linear regret. In contrast, fixed positive Gaussian value noise yields a joint regret--hard-violation lower bound of $Ω(\min\{σ,1\}T/\log^2 T)$, even with exact gradients in one dimension. The hard-violation construction matches arbitrarily many moments while preserving a feasible-endpoint gap and constant endpoint probabilities. Together, the bounds separate uncertainty about hard feasibility from learnable signed budgets. Deterministic restarts remove the horizon input without changing either upper rate.

发表机构

  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑