arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13547cs.LGstat.ML

面向多臂老虎机中遗忘对手的最优切换遗憾

Toward Optimal Switching Regret for Multi-Armed Bandits with Oblivious Adversary

Mengxiao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种结合固定份额学习器与二元区间子程序的算法,在遗忘对手下对任意未知切换次数S实现最优切换遗憾,解决了开放问题。

中文摘要 AI 辅助

我们研究对抗性多臂老虎机中的切换遗憾,其中学习器与一个最多改变$S$次的臂序列竞争。当$S$已知时,可以获得$\tilde{\mathcal{O}}(\sqrt{(S+1)KT})$的最优期望遗憾[Auer等人,2002]。然而,当$S$未知时,Marinov和Zimmert[2021]表明,在自适应对手下这一保证不可能实现。在本文中,我们证明了一个单一算法对遗忘对手针对每个$S$都能实现$\tilde{\mathcal{O}}(\sqrt{(S+1)KT})$的期望遗憾,解决了Auer等人[2019b]的一个开放问题。我们的算法结合了一个以较小学习率初始化的固定份额学习器,以及使用随机学习率和隐式探索来搜索局部改进的二元区间子程序。重要的是,非均匀先验倾向于跟随主学习器,从而保持维护许多子程序的成本较小。当子程序相对于主学习器积累了足够的改进时,其学习率会加倍,从而适应未知的比较器切换次数$S$。

英文摘要

We study switching regret in adversarial multi-armed bandits, where the learner competes with an arm sequence that changes at most $S$ times. When $S$ is known, an optimal expected regret of $\widetilde{\mathcal{O}}(\sqrt{(S+1)KT})$ is obtainable [Auer et al., 2002]. However, when $S$ is unknown, Marinov and Zimmert [2021] show that this guarantee is impossible under an adaptive adversary. In this paper, we show that a single algorithm achieves $\widetilde{\mathcal{O}}(\sqrt{(S+1)KT})$ expected regret for every $S$ against an oblivious adversary, resolving an open problem of Auer et al. [2019b]. Our algorithm combines a fixed-share learner initialized with a small learning rate and dyadic-interval subroutines that search for local improvements using randomized learning rates and implicit exploration. Importantly, a non-uniform prior favors following the main learner, keeping the cost of maintaining many subroutines small. When the subroutines accumulate sufficient improvement over the main learner, its learning rate doubles, allowing adaptation to the unknown number of comparator switches $S$.

发表机构

  • University of Iowa(爱荷华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑