发表机构
Politecnico di Milano; University of Ottawa; University of Bristol(米兰理工大学; 渥太华大学; 布里斯托尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对对抗性m-集多臂老虎机问题,提出一种高效算法,其遗憾界与EXP3-KW算法相当但仅需多项式时间,解决了相关开放问题。
AI 中文摘要
我们研究带有m-集动作的对抗性组合多臂老虎机问题,每一轮中学习器从d个物品中选择m个,仅观测所选物品的聚合损失。由此产生的动作集包含K=组合数(d,m)个元素,因此可能呈指数级庞大。不过,每个动作的损失由相同的d维物品损失向量决定。我们提出一种计算高效的算法,该算法利用此结构而无需显式枚举动作集。针对自适应非预知对手,该算法以至少1-δ的概率保证,相对于最优固定动作的遗憾为R_T=O(√(dT log(K/δ)))。这一遗憾界与Zimmert和Lattimore提出的有限动作EXP3-KW算法的高概率遗憾界匹配,而EXP3-KW算法的直接实现可能需要指数级空间。相反,我们的算法用d个参数表示每个采样分布,以多项式时间运行且无需枚举动作集,从而解决了Maiti等人提出的开放问题。
英文摘要
We study adversarial combinatorial bandits with $m$-set actions, where at each round the learner selects $m$ out of $d$ items and observes only the aggregate loss of the selected items. The resulting action set contains $K=\binom{d}{m}$ elements and can therefore be exponentially large. Nevertheless, the loss of every action is determined by the same $d$-dimensional vector of item losses. We propose a computationally efficient algorithm that exploits this structure without explicitly enumerating the action set. Against adaptive non-anticipating adversaries, it guarantees, with probability at least $1-δ$, regret against the best fixed action of \[ R_T = O\left(\sqrt{dT\log(K/δ)}\right). \] This matches the high-probability regret bound of the finite-action EXP3-KW algorithm of Zimmert and Lattimore, whose direct implementation may require exponential space. Our algorithm instead represents each sampling distribution with $d$ parameters and runs in polynomial time without enumerating the action set. Thus, it resolves the open problem posed by Maiti et al. We complement this upper bound with a matching high-probability lower bound. For all sufficiently small $δ$, every randomized policy admits a deterministic adaptive non-anticipating adversary for which, with probability at least $δ$, \[ R_T = Ω\left(\sqrt{dT\log(K/δ)}\right). \] Thus, the rate is minimax optimal up to universal constants in this regime. In particular, setting $m=1$ proves that the $\log K$ for ordinary $K$-armed bandits against adaptive non-anticipating adversaries is unavoidable, closing the remaining $\sqrt{\log K}$ gap between confidence-tuned upper and lower bounds left by Gerchinovitz and Lattimore.