发表机构
Department of Industrial Engineering and Operations Research, University of California, Berkeley(加州大学伯克利分校工业工程与运筹学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出以贝尔曼为中心的RSM-LinUCB算法,用于带记忆的线性老虎机,通过新颖的遗憾分解将记忆误差与奖励估计误差分离,实现近最优遗憾界,并扩展到广义线性奖励,实验验证优于基线。
AI 中文摘要
我们研究带记忆的线性老虎机问题,其中过去的动作通过一个任意已知、有界的矩阵值记忆映射引发内生的非平稳性。为了在考虑记忆动态的同时权衡探索与利用,我们开发了RSM-LinUCB,这是一种以贝尔曼为中心的算法,其学习方式如同线性老虎机,而规划方式如同强化学习。这一设计实现了一种新颖的遗憾分解,将记忆引发的误差与学习者轨迹上的累积奖励估计误差分离开来。我们证明了一个高概率遗憾界,为$\widetilde O\big(dRS(M+1)+\sigma d\sqrt T\big)$,其中$T$是学习时域,$d$是参数维度,$M$是记忆长度,$R$和$S$分别界定记忆映射算子范数和奖励参数范数,$\sigma$是次高斯噪声尺度。我们的结果表明,先前界中记忆与时域的乘法耦合并非内在固有:记忆仅贡献一个加性成本,直至对数因子。我们还证明了一个匹配的极小极大下界,确立了近最优性。我们进一步将算法扩展到广义线性奖励,在保持这种分离的同时,实现近最优的记忆和主要的统计依赖性。在合成实例以及半合成的KV和语义缓存任务上的数值实验中,我们的算法优于基线。
英文摘要
We study linear bandits with memory, where past actions induce endogenous nonstationarity through an arbitrary known, bounded matrix-valued memory map. To trade off exploration and exploitation while accounting for the memory dynamics, we develop RSM-LinUCB, a Bellman-centric algorithm that learns as in linear bandits and plans as in reinforcement learning. This design admits a novel regret decomposition which separates the memory-induced error from the cumulative reward estimation error along the learner's trajectory. We prove a high-probability regret bound of $\widetilde O\big(dRS(M+1)+σd\sqrt T\big)$, where $T$ is the learning horizon, $d$ is the parameter dimension, $M$ is the memory length, $R$ and $S$ bound the memory-map operator norm and reward-parameter norm, respectively, and $σ$ is the sub-Gaussian noise scale. Our results reveal that the multiplicative memory-horizon coupling in prior bounds is not intrinsic: memory only contributes an additive cost, up to logarithmic factors. We also prove a matching minimax lower bound, establishing near-optimality. We further extend the algorithm to generalized linear rewards, preserving this separation with near-optimal memory and leading statistical dependence. Our algorithms outperform the baselines in numerical experiments on synthetic instances and semi-synthetic KV- and semantic-cache tasks.