arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LinUCB的采样分配:小间隔机制下的最优设计极限

Sampling Allocation of LinUCB: Optimal Design Limits in the Small-Gap Regime

Yujie Liu, Vincent Y. F. Tan, Yunbei Xu

arXiv 2610.06213首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究LinUCB在小间隔机制下的采样分配,证明其经验采样分布收敛到D-最优设计,并据此细化遗憾分析、建立估计器的中心极限定理。

AI 中文摘要

我们研究了LinUCB在小间隔机制下的采样分配问题,其中奖励间隔在决策时间范围$n$内至多为$n^{-1/2}$量级。这种缩放捕捉了支撑最坏情况遗憾下界困难实例的特性,对于这些实例,LinUCB已知在$n$的对数因子范围内接近最优。利用平均场视角,我们通过经验采样分布来表征这种分配,这是一个宏观对象,它平均了时间范围内自适应决策的影响,并确定了其当$n\to\infty$时的极限。我们证明在该机制下,由LinUCB诱导的经验采样分布收敛到D-最优设计集合。这一核心结果揭示,在小间隔机制下,LinUCB不仅实现了接近最优的极小极大遗憾,而且以一种对学习奖励参数渐近有效的方式分配样本,从而将遗憾驱动的在线学习与信息高效的实验设计联系起来。基于最优设计极限,我们获得了两个有用的推论。首先,我们通过刻画其极限中的前导阶常数,细化了LinUCB在小间隔机制下的渐近遗憾分析。其次,我们表明,尽管LinUCB采用自适应采样策略,正则化最小二乘估计器在小间隔机制下满足中心极限型定理,从而能够对奖励参数进行有效的统计推断。

英文摘要

We study the sampling allocation of LinUCB in the small-gap regime, where the reward gaps are of order at most $n^{-1/2}$ over the decision horizon $n$. This scaling captures the hard instances underlying worst-case regret lower bounds, for which LinUCB is known to be near optimal up to logarithmic factors in $n$. Using a mean-field perspective, we characterize this allocation through the empirical sampling distribution, a macroscopic object that averages the effect of adaptive decisions over the horizon, and identify its limit as $n\to\infty$. We establish that in this regime, the empirical sampling distribution induced by LinUCB converges to the set of D-optimal designs. This central result reveals that, in the small-gap regime, LinUCB not only achieves near optimal minimax regret but also allocates samples in a way that is asymptotically efficient for learning the reward parameter, thereby connecting regret-driven online learning with information-efficient experimental design. Building on the optimal design limit, we obtain two useful consequences. First, we refine the asymptotic regret analysis of LinUCB in the small-gap regime by characterizing its leading-order constant in the limit. Second, we show that, despite LinUCB's adaptive sampling strategy, the regularized least-squares estimator satisfies a central-limit-type theorem in the small-gap regime, thereby enabling valid statistical inference for the reward parameter.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑