发表机构
Institute for Interdisciplinary Information Sciences Tsinghua University(清华大学交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对任意反馈延迟下的Lipschitz多臂老虎机,分别为随机和对抗奖励设置提出消除算法与EXP3算法,所得后悔界与无延迟时匹配,刻画了反馈延迟的额外影响。
AI 中文摘要
Lipschitz多臂老虎机问题将传统多臂老虎机框架扩展到连续动作空间,假设奖励函数满足Lipschitz条件。本研究探讨具有任意反馈延迟的Lipschitz多臂老虎机,即执行动作后无法立即获得奖励信号,而是经过任意选择的延迟后才能收到。我们考虑随机奖励和对抗奖励两种设置,分别提出一种基于消除(elimination)的算法和一种基于EXP3的算法。对于两种设置,我们的算法在时间范围T和总延迟D下均达到后悔界$\tilde{O}\bigl(T^{\frac{d_z+1}{d_z+2}}+\bigl(\tilde{\bigl(}\bigr)\bigr)\bigr)$,其中两种设置的主要区别在于缩放维度$d_z$的定义。我们的界与无延迟时Lipschitz多臂老虎机的现有后悔保证相匹配,并刻画了反馈延迟引入的额外$\tilde{O}(\bigl(\bigr)\bigr)$影响。
英文摘要
The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz condition. This work investigates Lipschitz bandits under arbitrary feedback delays, where reward signals are not received immediately upon taking an action but after an arbitrarily chosen delay. We consider both stochastic and adversarial reward settings, proposing an elimination-based algorithm and an EXP3-based algorithm, respectively. For both settings, our algorithms achieve a regret bound of $\tilde{O}\left(T^{\frac{d_z+1}{d_z+2}}+\sqrt{D}\right)$ over a time horizon $T$ with total delay $D$, where the main difference between settings lies in the definition of the zooming dimension $d_z$. Our bounds match existing delay-free regret guarantees for Lipschitz bandits and characterize the additional $\tilde{O}(\sqrt{D})$ impact introduced by feedback delays.
Comments10 pages of main contents, 26 pages in total