arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

精确滑动窗口约束下的线性赌博机

Linear Bandits under Exact Sliding-Window Constraints

Seyed Mohammad Hadi Hosseini, Yasin Abbasi-Yadkori, Sattar Vakili

arXiv 2610.08745首次发表:更新:

AI 中文总结

本文研究精确滑动窗口约束下的线性赌博机,提出转移直径与历史状态直径,开发稀有切换OFUL算法,实现亚线性遗憾并保持精确可行性。

AI 中文摘要

我们研究了精确滑动窗口约束下的线性赌博机问题,其中每个连续的动作块必须属于一个规定的可行集合。在离线设置中,即奖励函数已知的情况下,我们证明了当 $w\mid T$ 时,凸性和循环移位不变性使得平稳解是最优的,否则在加性 $O(w)$ 的差距内也是最优的。在在线设置中,我们表明仅凭几何结构不足以进行学习,亚线性遗憾可能无法实现。我们引入了一个转移直径 $\tau$ 来量化可行的可达性,并开发了一种稀有切换的 OFUL 算法,其遗憾为 $\widetilde{O}(d\sqrt{T}+\tau d+w)$,相对于离线最优可行轨迹。最后,我们移除循环不变性并考虑一般的滑动窗口约束,其中最优行为可能是非平稳的。我们将最近的动作历史表示为有限记忆控制问题的状态,并引入一个历史状态直径 $D$ 来度量可行历史之间的可行通信。结合乐观剩余视野规划与稀有策略更新,我们获得了 $\widetilde{O}(d\sqrt{T}+dD+w)$ 的遗憾界。我们在真实世界和合成基准上评估了我们的方法,表明它在保持精确可行性的同时,实现了与基线相当的奖励和遗憾,且策略更新次数显著减少。

英文摘要

We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity and cyclic-shift invariance make a stationary solution optimal when $w\mid T$ and within an additive $O(w)$ gap otherwise. In the online setting, we show that geometric structure alone is insufficient for learning, and sublinear regret can be impossible. We introduce a transition diameter $τ$ that quantifies feasible reachability and develop a rare-switching OFUL algorithm with regret $\widetilde{O}(d\sqrt{T}+τd+w)$ against the offline-optimal feasible trajectory. Finally, we remove cyclic invariance and consider general sliding-window constraints, where optimal behavior may be non-stationary. We represent recent action history as the state of a finite-memory control problem and introduce a history-state diameter $D$ that measures feasible communication between viable histories. Combining optimistic remaining-horizon planning with rare policy updates, we obtain a regret bound of $\widetilde{O}(d\sqrt{T}+dD+w)$. We evaluate our approach on real-world and synthetic benchmarks, showing that it maintains exact feasibility while achieving reward and regret comparable to baselines with substantially fewer policy updates.

Comments53 pages, including supplementary material; 8 figures and 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑