发表机构
Abstract Math Institute(抽象数学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究高维高斯线性老虎机,推导汤普森采样与贪心选择的精确贝叶斯遗憾,证明贪心策略渐近最优而汤普森采样遗憾更大。
AI 中文摘要
我们研究贝叶斯线性多臂老虎机问题,其中参数服从各向同性高斯分布,候选臂独立同高斯分布,奖励噪声为高斯噪声,且时间范围与维度成比例。归一化后验不确定性具有一个显式极限,该极限对所有因果策略一致成立。高斯后验恒等式进而确定极限参数重叠,而无需假设自适应递归的封闭性。这些结果给出了汤普森采样、后验均值贪心选择以及一族缩放后验采样协方差的策略的精确遗憾曲线。归一化实现的累积遗憾在紧致比例时间区间上一致地以L1范数收敛。一个策略一致的界确定了极限最优贝叶斯遗憾,并证明后验均值贪心选择能够达到该界。汤普森采样产生严格更大的主导遗憾;其相对于贪心选择的瞬时遗憾比率介于1和2之间,并在长比例时间范围趋近于2。闭式累积曲线还揭示了在噪声消失极限下的不同比较。最后,瞬时遗憾收敛于一个非退化的高斯决策损失分布,而非其均值。该分析将老虎机策略获取的信息量与其利用该信息所做决策的质量分离开来。
英文摘要
We study Bayesian linear bandits with an isotropic Gaussian parameter, independent Gaussian candidate arms, and Gaussian reward noise when the horizon is proportional to the dimension. The normalized posterior uncertainty has an explicit limit that is uniform over all causal policies. Gaussian posterior identities then determine the limiting parameter overlaps without an assumed closure of the adaptive recursion. These results yield exact regret curves for Thompson sampling, posterior-mean greedy selection, and a family of policies that scale the posterior sampling covariance. The normalized realized cumulative regret converges in L1, uniformly on compact proportional-time intervals. A policy-uniform lower bound identifies the limiting optimal Bayes regret and proves that posterior-mean greedy selection attains it. Thompson sampling incurs a strictly larger leading regret; its instantaneous regret ratio relative to greedy selection lies between one and two and approaches two at long proportional horizons. Closed-form cumulative curves also identify a different comparison in the vanishing-noise limit. Finally, the instantaneous regret converges to a nondegenerate Gaussian decision-loss distribution, rather than to its mean. The analysis separates the amount of information acquired by a bandit policy from the quality of the decisions made using that information.
Comments17 Pages