发表机构
Department of Computer Science, University of Colorado, Boulder; INRIA, Paris(科罗拉多大学博尔德分校计算机科学系; 法国国家信息与自动化研究所巴黎分部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究动态图上随机多臂强盗问题,基于滑动窗口混合识别结构条件,分析局部探索然后提交算法,建立次线性期望遗憾,还包括奖励感知策略及相关定理。
AI 中文摘要
我们研究动态图上的随机多臂强盗问题,其中臂对应具有时变边的网络顶点。在此设置中,学习者限于局部移动,每轮仅选择其当前节点或直接邻居。这种约束使最佳臂识别与利用解耦:即使识别出最优臂,学习者可能仍无法通过不断演变的拓扑结构到达它。我们基于滑动窗口混合识别了一个与过程无关的结构条件,确保图的固有游走对于探索和导航都保持稳定。在此框架下,我们分析了一类局部探索然后提交算法并建立了次线性期望遗憾。我们的框架包括一个奖励感知策略,为此我们证明了一个最坏情况安全定理和一个单独的性能增益定理。
英文摘要
We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In this setting, the learner is restricted to local movement, selecting only its current node or an immediate neighbor at each round. This constraint decouples best-arm identification from exploitation: even after the optimal arm is identified, the learner may remain unable to reach it through the evolving topology. We identify a process-agnostic structural condition, based on sliding-window mixing, that ensures the graph's intrinsic walk remains stable for both exploration and navigation. Under this regime, we analyze a family of local explore-then-commit algorithms and establish sublinear expected regret. Our framework includes a reward-aware strategy, for which we prove a worst-case safety theorem and a separate performance gain theorem.