arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有部分观测动作的随机线性博弈

Stochastic Linear Bandits with Partially Observed Actions

Gautam Dasarathy, Vineet Gattani, Lalit Jain

arXiv 2607.08971首次发表:更新:

发表机构

Arizona State University; GE Vernova; Google(亚利桑那州立大学; 通用电气 Vernova; 谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究具有部分观测动作的随机线性博弈问题,提出TOFU-POV算法,通过估计潜在动作子空间等操作,使遗憾为$\sqrt{T}$且与内在维度相关,还设计了秩自适应算法,实验证明该算法能改进自然基线。

AI 中文摘要

随机线性博弈是顺序决策的核心范式,动作表示为向量,奖励是线性的。本文研究其部分观测变体,即学习智能体每次动作只能看到坐标的随机子集,这在推荐和医疗等场景自然出现。一般情况下,次线性遗憾在信息理论上不可能。但当动作向量具有低内在维度时可克服此障碍。提出算法TOFU-POV,用掩码动作估计潜在动作子空间,用逐轮冻结表示插补当前动作,并在低维坐标中运行OFUL。理论表明TOFU-POV的遗憾为$\sqrt{T}$,与内在动作子空间维度而非环境维度相关,并量化了这些量与缺失、决策集大小和子空间条件之间的相互作用。还设计了无需内在维度知识的秩自适应算法。通过合成和真实数据实验支持理论,表明TOFU-POV在该挑战性问题上可大幅改进自然基线。

英文摘要

The stochastic linear bandit, where actions are represented as vectors and rewards are linear, is a central paradigm for sequential decision making. We study a partially observed variant of this problem in which the learning agent only sees a random subset of coordinates for each action. Such partial observability arises naturally in settings like recommendation and healthcare, where full action descriptions can be expensive or even impossible to obtain. In general, this makes sublinear regret information-theoretically impossible. However, we show that this barrier can be overcome when the action vectors have low intrinsic dimension. We propose an algorithm, TOFU-POV, that estimates the latent action subspace using the masked actions, imputes current actions using an epoch-wise frozen representation, and runs OFUL in the resulting low-dimensional coordinates. Our theory shows that TOFU-POV enjoys a $\sqrt{T}$ regret that scales with the intrinsic action subspace dimension as opposed to the ambient dimension and quantifies the interaction between these quantities and the missingness, decision set size, and subspace conditioning. We also devise a rank-adaptive algorithm that does not require the knowledge of the intrinsic dimension. We complement these guarantees with a lower bound based on a novel product construction that separates usual reward-learning uncertainty from a missingness-dependent cost intrinsic to partial observation. Synthetic and real data experiments support our theory and show that TOFU-POV can substantially improve upon natural baselines in this challenging problem.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑