arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从具有强盗反馈的动态图上的局部游走中学习

Learning from Local Walks on Dynamic Graphs with Bandit Feedback

Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni, Lijun Chen

arXiv 2607.10571首次发表:更新:

发表机构

Department of Computer Science, University of Colorado, Boulder; INRIA, Paris(科罗拉多大学博尔德分校计算机科学系; 法国国家信息与自动化研究所巴黎分部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究动态图上随机多臂强盗问题,基于滑动窗口混合识别结构条件,分析局部探索然后提交算法,建立次线性期望遗憾,还包括奖励感知策略及相关定理。

AI 中文摘要

我们研究动态图上的随机多臂强盗问题,其中臂对应具有时变边的网络顶点。在此设置中,学习者限于局部移动,每轮仅选择其当前节点或直接邻居。这种约束使最佳臂识别与利用解耦:即使识别出最优臂,学习者可能仍无法通过不断演变的拓扑结构到达它。我们基于滑动窗口混合识别了一个与过程无关的结构条件,确保图的固有游走对于探索和导航都保持稳定。在此框架下,我们分析了一类局部探索然后提交算法并建立了次线性期望遗憾。我们的框架包括一个奖励感知策略,为此我们证明了一个最坏情况安全定理和一个单独的性能增益定理。

英文摘要

We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In this setting, the learner is restricted to local movement, selecting only its current node or an immediate neighbor at each round. This constraint decouples best-arm identification from exploitation: even after the optimal arm is identified, the learner may remain unable to reach it through the evolving topology. We identify a process-agnostic structural condition, based on sliding-window mixing, that ensures the graph's intrinsic walk remains stable for both exploration and navigation. Under this regime, we analyze a family of local explore-then-commit algorithms and establish sublinear expected regret. Our framework includes a reward-aware strategy, for which we prove a worst-case safety theorem and a separate performance gain theorem.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑