arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.25029cs.LG

在线凸优化中具有两点老虎机反馈的最优高概率懊悔

Logarithmic High-Probability Regret for Online Convex Optimization with Two-Point Bandit Feedback

Haishan Ye

AI总结:

本文研究在线凸优化问题,在两点老虎机反馈下,提出首个针对强凸损失函数的高概率懊悔界,结果在时间跨度和维度上均最优。

AI中文摘要:

我们考虑在线凸优化(OCO)问题,其中玩家试图最小化一系列由对手生成的凸损失函数,但只能观察每个函数在两点处的值。尽管两点反馈允许梯度估计,但实现强凸函数的紧高概率懊悔界仍是一个开放问题,如Agarwal等人所指出的。主要挑战在于老虎机梯度估计的厚尾性质,使得标准集中分析困难。本文解决了这一开放问题,提供了首个高概率懊悔界$O(d(\log T + \log(1/δ))/μ)$,针对μ-强凸损失函数。结果在时间跨度$T$和维度$d$上均最优。

英文摘要:

We study online convex optimization (OCO) with two-point bandit feedback against a non-anticipating adaptive adversary. In this setting, a learner competes with an adversarial sequence of convex losses while observing each loss only through two function evaluations. For strongly convex losses, Agarwal, Dekel, and Xiao~\citeyearpar{agarwal2010optimal} proved a comparator-wise logarithmic regret bound in expectation. Consequently, by minimizing outside the probability space, their result yields a pseudo-regret guarantee of the form $\EB A_T-\min_{x\in\mathcal K}\EB L_T(x)$, where $A_T$ is the algorithm's two-query cumulative loss and $L_T(x)$ is the comparator's cumulative loss. They asked whether a logarithmic high-probability guarantee is achievable in the same two-point strongly convex setting. Our main theorem provides the corresponding fixed-comparator high-probability statement: for any comparator $x\in\mathcal K$ fixed independently of the algorithmic random directions, the standard two-point projected gradient method guarantees, with probability at least $1-δ$, a two-query regret bound of order \[ O\left(\frac{dG^2}μ\left(\log T+\log(1/δ)\right)+dGD\log(1/δ)+G\log T\left(1+\frac{D}{r}\right)\right). \] At the comparator-wise level, our leading horizon-dependent term is linear in $d$, compared with the $d^2$-type term in the original analysis of Agarwal, Dekel, and Xiao. The key ingredient is a high-confidence analysis that simultaneously absorbs the martingale error into strong convexity and preserves the linear-in-dimension estimator control of the two-point method. A deterministic covering argument then yields a realized full-comparator guarantee against $\min_{x\in\mathcal K}L_T(x)$, preserving logarithmic dependence on $T$ at the cost of the standard covering-number factor.

↑