arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有潜在线性动力学的非平稳赌博机的块乐观算法

Block Optimism for Nonstationary Bandits with Latent Linear Dynamics

Taehyun Hwang, Hyunjun Choi, Heesang Ann, Min-hwan Oh

arXiv 2610.00911首次发表:更新:

发表机构

Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对具有潜在线性动力学的非平稳赌博机,提出基于循环近似和UCB的块级乐观算法,将遗憾从$\tilde{O}(T^{2/3})$改进至$\tilde{O}(\sqrt T)$。

AI 中文摘要

我们研究了一个具有潜在线性动力学的内生非平稳随机赌博机问题,其中动作既影响即时奖励,也影响未观测潜在状态的未来演化。奖励是当前动作与潜在状态的双线性函数,从而产生依赖于历史的奖励和非平凡的长期规划问题。现有的先探索后提交方法通过均匀探索来估计潜在动力学,然后提交优化的开环动作序列,实现了$\tilde{O}(T^{2/3})$的遗憾。我们表明,通过自适应块级乐观可以改进这一速率。我们的关键步骤是一个循环近似:在稳定动力学下,无限记忆的奖励过程可以被截断,开环基准可以通过优化有限记忆的块级代理来近似。基于这一归约,我们提出了一种基于UCB的块算法,该算法维护截断动力学参数的置信集,并乐观地选择块。我们证明了遗憾界为$\tilde{O}(\sqrt T)$,显著优于同一模型的先前$\tilde{O}(T^{2/3})$保证。据我们所知,这是针对具有双线性奖励观测和开环动作序列基准的潜在线性动力学赌博机的第一个$\tilde{O}(\sqrt T)$遗憾保证。

英文摘要

We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon planning problem. The existing explore-then-commit approach achieves $\tilde{O}(T^{2/3})$ regret by uniformly exploring to estimate the latent dynamics and then committing to an optimized open-loop action sequence. We show that this rate can be improved via adaptive block-level optimism. Our key step is a cyclic approximation: under stable dynamics, the infinite-memory reward process can be truncated, and the open-loop benchmark can be approximated by optimizing a finite-memory block-level proxy. Building on this reduction, we propose a UCB-based block algorithm that maintains confidence sets for the truncated dynamics parameters and selects blocks optimistically. We prove a regret bound of order $\tilde{O}(\sqrt T)$, significantly improving over the previous $\tilde{O}(T^{2/3})$ guarantee for the same model. To the best of our knowledge, this is the first $\tilde{O}(\sqrt T)$ regret guarantee for latent linear-dynamics bandits with bilinear reward observations and an open-loop action-sequence benchmark.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑