arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单副本量子老虎机是经典的:精确谱坍缩

The Single-Copy Quantum Bandit Is Classical: An Exact Spectral Collapse

Siu Hin Ng

arXiv 2609.40339首次发表:更新:

发表机构

Academia Sinica(中央研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究证明单副本量子老虎机中谱测量加经典KL-UCB达到渐近实例最优,其遗憾下界由经典Burnetas-Katehakis泛函决定,并指出进一步改进需超越单副本测量类。

AI 中文摘要

我们研究多臂老虎机问题,其中每个臂提供未知的量子态,其平均奖励由已知效应$F$定义。每个新副本可通过任意自适应选择的POVM进行测量。我们证明了一个精确坍缩:从臂态到其混淆奖励半空间的测量相对熵等于$F$谱统计的经典Burnetas-Katehakis泛函,并且$F$的谱测量与一个由矩阵对数导数构建的显式最不利态一起,构成了底层测量博弈的鞍点。因此,在所有一致的单副本策略中,谱测量后接经典KL-UCB算法在渐近意义上是实例最优的,适用于每个有限维度以及在所述非退化假设下的每个奖励效应:自适应和随机测量设计无法改进前导对数遗憾系数,存储副本以供后续单副本测量也无法改进。一个自包含的有限时间界涵盖了一般效应和边界分布。任何进一步的改进必须来自单副本类之外的测量。对于具有任意量子存储器的一致策略,一个摊销相对熵论证给出了与Umegaki半空间散度的逆命题,并且两个每臂遗憾下界$1/K_{\mathrm{inf}}$和$1/D_{\mathrm{inf}}$当且仅当臂与$F$对易时重合。Umegaki下界是否可达涉及一个复合量子Stein问题以及一个到自适应遗憾的单独归约;我们将其作为开放问题提出,并附有两副本数值证据。

英文摘要

We study multi-armed bandits whose arms supply unknown quantum states and whose mean rewards are defined by a known effect $F$. Each fresh copy may be measured by an arbitrary, adaptively chosen POVM. We prove an exact collapse: the measured relative entropy from an arm state to its confusing reward half-space equals the classical Burnetas-Katehakis functional of $F$'s spectral statistics, and the spectral measurement of $F$ together with an explicit least-favorable state, built from the derivative of the matrix logarithm, forms a saddle point of the underlying measurement game. Consequently, spectral measurement followed by classical KL-UCB is asymptotically instance-optimal among all consistent single-copy policies, in every finite dimension and for every reward effect under the stated nondegeneracy assumptions: adaptive and randomized measurement design cannot improve the leading logarithmic regret coefficient, and neither can storing copies for later single-copy measurement. A self-contained finite-time bound covers general effects and boundary distributions. Any further improvement must come from measurements outside the single-copy class. For consistent policies with arbitrary quantum memory, an amortized relative-entropy argument gives a converse with the Umegaki half-space divergence, and the two per-arm regret floors, $1/K_{\mathrm{inf}}$ and $1/D_{\mathrm{inf}}$, coincide if and only if the arm commutes with $F$. Whether the Umegaki floor is attainable involves a composite quantum Stein problem and a separate reduction to adaptive regret; we pose it as an open problem, with two-copy numerical evidence.

Comments9 pages, 4 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑