arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15996cs.LG

面向对抗性多臂老虎机的最优二阶路径长度保证

Toward Optimal Second-Order Path-Length Guarantee for Adversarial Multi-Armed Bandits

  • University of Iowa(爱荷华大学)

机构由 AI 辅助整理,请以论文原文为准。

Mengxiao Zhang

中文总结 AI 辅助

该研究针对对抗性多臂老虎机的二阶路径长度后悔问题,证明Bubeck等人的算法在已知二阶路径长度时可达到匹配下界的期望后悔,还通过自适应重启方案消除了对该参数的先验知识要求。

中文摘要 AI 辅助

我们研究对抗性K臂老虎机在对抗无感知损失序列时的二阶路径长度后悔。Bubeck等人[2019]设计了一种算法,其后悔达到\widetilde{\mathcal{O}}(K+\sqrt{KQ_{\infty,1}}),其中Q_{\infty,1}是一阶路径长度,该研究留下一个问题:在老虎机反馈下是否能达到\widetilde{\mathcal{O}}(\text{poly}(K)\sqrt{1+Q_{\infty,2}})的后悔,Q_{\infty,2}是二阶路径长度。出人意料的是,我们正面解决了该问题,通过更复杂的分析表明,当Q_{\infty,2}已知时,Bubeck等人[2019]的同一种算法能达到\mathcal{O}\left(K\log(KT)+\sqrt{K\log(KT)\bigl(1+Q_{\infty,2}\bigr)}\right)的期望后悔,T是时间范围,该结果在对数因子和加项上与\Omega(\sqrt{KQ_{\infty,2}})的下界匹配。我们还通过一种路径长度估计量增量有界的自适应重启方案,消除了对Q_{\infty,2}的先验知识要求。

英文摘要

We study second-order path-length regret in adversarial $K$-armed bandits against oblivious loss sequences. Bubeck et al. [2019] designed an algorithm that achieves $\widetilde{\mathcal{O}}(K+\sqrt{KQ_{\infty,1}})$ regret, where $Q_{\infty,1}$ is the first-order path length, and left open whether $\widetilde{\mathcal{O}}(\text{poly}(K)\sqrt{1+Q_{\infty,2}})$ regret is achievable under bandit feedback, where $Q_{\infty,2}$ is the second-order path length. Somewhat surprisingly, we resolve this question positively by showing that with a more involved analysis, the exact same algorithm of Bubeck et al. [2019] achieves $\mathcal{O}\left(K\log(KT)+\sqrt{K\log(KT)\bigl(1+Q_{\infty,2}\bigr)}\right)$ expected regret when $Q_{\infty,2}$ is known, where $T$ is the horizon. This matches the $Ω(\sqrt{KQ_{\infty,2}})$ lower bound up to logarithmic factors and additive terms. We further remove the knowledge of $Q_{\infty,2}$ using an adaptive restart scheme whose path-length estimator has uniformly bounded increments.

↑