arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

球形动作集上线性赌博机的一种统一乐观无关框架

A Unified Optimism-Agnostic Framework for Linear Bandits over Spherical Action Sets

Arda Güçlü, Subhonmesh Bose, John R. Birge

arXiv 2609.32149首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; University of Chicago Booth School of Business(伊利诺伊大学厄巴纳-香槟分校; 芝加哥大学布斯商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一个统一框架,通过设计矩阵最小特征值增长和动作集中性条件,证明UCB和TS变体在球形动作集上达到最优遗憾率,连接参数估计质量与遗憾累积。

AI 中文摘要

线性赌博机模型描述了序贯决策问题,其中噪声奖励是决策变量的线性函数,智能体必须同时学习控制平均奖励的未知参数,同时随时间最大化(期望)奖励。两个著名的算法家族——上置信界(UCB)和汤普森采样(TS)——在探索(以估计该参数)和利用(利用其知识)之间取得平衡。该参数估计的质量取决于设计矩阵的特征值。在本文中,我们首先表明,如果从探索中获得的推断质量(以设计矩阵的最小特征值编码)随时间$t$增长$\gtrsim \sqrt{t}$,同时动作保持足够集中以进行利用,那么对于球形动作集,算法在时间范围$T$内产生最优的高概率$\mathcal{O}(\sqrt{T}\log T)$遗憾率。该分析是算法无关的,并遵循经典基于乐观的椭圆势参数进行遗憾分析的替代途径。然后,我们说明UCB和TS的变体满足推断和集中性质,并因此享有最优遗憾率。实际上,我们的结果提供了一个模块化框架,可用于分析线性赌博机算法,并明确地将参数估计质量与最优遗憾累积联系起来。

英文摘要

Linear bandits model sequential decision-making problems with noisy rewards that are linear in the decision variable, where an agent must simultaneously learn about an unknown parameter that governs the mean rewards, while maximizing (expected) rewards over time. Two prominent algorithmic families--upper confidence bound (UCB) and Thompson sampling (TS)--achieve a balance of exploration (to estimate said parameter) and exploitation (utilization of knowledge about it) across time. The quality of estimation of that parameter depends on the eigenvalues of a design matrix. In this paper, we begin by showing that if the inference quality obtained from exploration, encoded in the minimum eigenvalue of the design matrix, grows $\gtrsim \sqrt{t}$ with time $t$, while actions remain sufficiently concentrated for exploitation, then an algorithm produces optimal high-probability $\mathcal{O}(\sqrt{T}\log T)$-regret rate over a time-horizon $T$ for spherical action sets. This analysis is algorithm-agnostic and follows an alternative route to the classical optimism-based elliptical-potential argument for regret analysis. Then, we illustrate that variants of UCB and TS satisfy the inference and concentration properties and in turn, enjoy optimal regret rate. In effect, our results provide a modular framework that can be used to analyze linear bandit algorithms and explicitly connect quality of parameter estimation to optimal regret accumulation.

Comments27 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑