发表机构
CWI; University of Twente; Indian Institute of Science (IISc); Leiden University(荷兰数学与计算机科学研究所; 特文特大学; 印度科学学院; 莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对均值漂移的随机多臂老虎机,提出重要性权重算法ISM,实现高概率最优臂识别,样本复杂度为K(σ²+U²)Δ_min⁻²ln(1/δ),并给出匹配下界。
AI 中文摘要
我们研究随机环境中带有一种新型对抗性扰动的最佳臂识别问题,我们将这种扰动称为“均值漂移”(Shifting Means)。经典情况下,$K$ 个臂的平均奖励在时间上是稳定的,而在均值漂移问题中,只有平均奖励之间的间隙 $\boldsymbol{\Delta}$ 是稳定的,而它们的共同偏移量可能每轮由对抗性方式决定。学习者的目标是在最小化样本复杂度的同时,以高概率识别出最佳臂(固定置信度设置)。处理漂移需要新的工具:我们证明,采用广义似然比检验(GLRT)停止规则的算法,包括流行的 Track-and-Stop,在时变漂移下会失败。相反,我们提出了均值漂移的重要性权重算法($\mathsf{ISM}$)。假设奖励均值以 $U$ 为界且为 $\sigma^2$-次高斯分布,我们证明 $\mathsf{ISM}$ 是 $\delta$-正确的,并且其样本复杂度界为 $K (\sigma^2 + U^2) \Delta_{\min}^{-2} \ln \frac{1}{\delta}$ 量级。我们还提出了一个匹配(至多常数因子)的最坏情况下的下界,并实证评估了我们的结果。
英文摘要
We study the best arm identification problem in a stochastic environment with a novel form of adversarial perturbations, which we coin Shifting Means. While classically the mean rewards of the $K$ arms are stable in time, in Shifting Means only the gaps $\boldsymbolΔ$ between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means ($\mathsf{ISM}$). Assuming means bounded by $U$ and $σ^2$-sub-Gaussian rewards, we show $\mathsf{ISM}$ to be $δ$-correct and to enjoy a sample complexity bound of order $K (σ^2 + U^2) Δ_{\min}^{-2} \ln \frac{1}δ$. We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.
CommentsAccepted at NeurIPS 2026