arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14319cs.LGquant-ph

量子多臂老虎机与线性老虎机:下界与算法

Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms

  • The Chinese University of Hong Kong(香港中文大学)
  • Xidian University(西安电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Maoli Liu, Zhuohua Li, John C. S. Lui

AI总结:

该研究针对量子多臂老虎机和量子线性老虎机,证明了新的极小极大下界,并提出了改进维度依赖的基于设计的消除算法,解决了Wan等人提出的相关问题。

AI中文摘要:

我们在Wan等人[2023]提出的模型中研究量子多臂老虎机(QMAB)和量子线性老虎机(QLB),其中学习者通过量子奖励神谕或其逆来查询每个臂或动作。现有研究针对K臂QMAB给出了时间范围T上遗憾值为O(K log T)的算法,针对d维QLB给出了遗憾值为O(d² polylog T)的算法,但这留下了两个未解决的问题:K log T的规模是否不可避免,以及d²的依赖关系是否可以改进。我们证明了QMAB的首个极小极大下界为Ω(K log(T/K)),有限动作QLB的极小极大下界为Ω(d log(T/d)),解决了Wan等人[2023]提出的“是否可实现与T无关的遗憾值”的问题。我们论证的核心是高置信度单臂量子测试下界,用于区分固定奖励均值与备选区间,该下界通过多项式方法和三角多项式的Remez型不等式证明。随后,老虎机到测试的归约将其提升为QMAB下界,而线性嵌入则给出有限动作QLB下界。作为下界的补充,我们为有限动作QLB提出了基于设计的消除算法:当动作集大小为poly(d)时,其遗憾值与d线性相关,改进了现有d²的依赖关系,且在多对数因子范围内与我们的下界匹配。该算法结合了低偏差低方差量子均值估计器与小支持G-最优设计,通过匹配设计权重的查询分配实现;当使用量子蒙特卡罗估计时,基于设计的消除将维度依赖从d²降低到d^(3/2),而低方差估计器使重构误差通过方差而非最坏情况绝对误差聚合,消除了剩余的√d因子。

英文摘要:

We study quantum multi-armed bandits (QMAB) and quantum linear bandits (QLB), where the learner queries each arm or action through a quantum reward oracle or its inverse. Prior work gives algorithms over horizon $T$ with regret $O(K\log T)$ for QMAB with $K$ arms and $O(d^2\operatorname{polylog} T)$ for $d$-dimensional QLB. This leaves open the optimal dependence on $K$ and $T$ and whether the dependence on $d$ can be further improved. In this work, we prove the first tight minimax regret bound of $Θ(K\log(1+T/K))$ for QMAB and the first lower bound of $Ω(d\log(1+T/d))$ for finite-action QLB, ruling out regret independent of $T$. Our lower bounds rely on a high-confidence single-arm quantum testing lower bound for distinguishing a fixed reward mean from an interval of alternatives. A bandit-to-testing reduction then lifts it to the QMAB lower bound, while a linear embedding gives the finite-action QLB lower bound. The matching QMAB upper bound is obtained using a tail bound for the Quantum Monte Carlo (QMC) estimator. For finite-action QLB, we propose a phased elimination algorithm that combines a low-bias low-variance quantum mean estimator with a small-support $G$-optimal design through a query allocation matched to the design weights. When the action set has size $\operatorname{poly}(d)$, its regret is nearly linear in $d$ and matches our lower bound up to polylogarithmic factors.

补充信息

↑