发表机构
Georgia Tech; UIC(佐治亚理工学院; 伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究强盗反馈版本的在线主成分分析,改进了遗憾值的上下界,弥合差距至\(r\sqrt{dT}\)。上界通过新算法实现,结合在线镜像下降与多尺度探索;下界通过构造自适应对手,将遗憾值下界估计转化为子空间估计问题,并讨论其与量子断层扫描联系。
AI 中文摘要
我们研究在线主成分分析的强盗反馈版本(强盗主成分分析):在每一轮\(t = 1,\dots,T\)中,对手选择一个\(d \times d\)对称增益矩阵\(G_t\),其谱在\([0,1]\)内且秩至多为\(r\);学习者同时选择一个单位向量\(w_t \in S^{d - 1}\)并接收奖励\(w_t^\top G_t w_t\)。学习者没有其他反馈,旨在最小化与事后最佳单位向量相比的遗憾值。这个问题由Kotlowski和Neu(2019)提出,他们给出了一个遗憾值为\(O(d\sqrt{rT \log T})\)的算法,并证明了下界为\(\Omega(r\sqrt{T/\log T})\)。我们改进了这两个界并基本弥合了它们之间的差距,在\(d\)和\(T\)的多对数因子范围内建立了阶为\(r\sqrt{dT}\)的极小极大遗憾值。上界由一种新颖算法实现,该算法将(实)密度矩阵谱面体上的在线镜像下降与多尺度探索方案相结合,其中具有不同谱大小的特征子空间以不同速率更新。对于下界,我们构造了一个自适应对手,它根据学习者的行动细化一个隐藏的大奖励子空间,使得在不估计子空间的情况下不可能有低遗憾值;因此,对遗憾值进行下界估计归结为研究出现的子空间估计问题。最后,我们讨论了强盗主成分分析与自适应测量量子断层扫描的联系。
英文摘要
We study the bandit-feedback version of online principal component analysis (Bandit PCA): in each round $t = 1,\dots,T$, the adversary selects a $d \times d$ symmetric gain matrix $G_t$ with spectrum in $[0,1]$ and rank at most $r$; the learner simultaneously selects a unit vector $w_t \in S^{d-1}$ and receives the reward $w_t^\top G_t w_t$. The learner receives no other feedback, and aims to minimize the regret against the best unit vector in hindsight. This problem was introduced by Kotlowski and Neu (2019), who gave an algorithm with regret $O(d\sqrt{rT \log T})$ and showed the lower bound of $Ω(r\sqrt{T/\log T})$. We improve upon both of these bounds and essentially bridge the gap between them, establishing the minimax regret of order $r\sqrt{dT}$ up to polylogarithmic factors in $d$ and $T$. The upper bound is attained by a novel algorithm, which combines online mirror descent on the spectrahedron of (real) density matrices with a multiscale exploration scheme in which the eigenspaces with different spectral magnitudes are updated at different rates. For the lower bound, we construct an adaptive adversary that refines a hidden large-reward subspace based on the learner's actions, in such a way that low regret is impossible without estimating the subspace; as a result, lower-bounding the regret reduces to studying the arising subspace estimation problem. Finally, we discuss connections of Bandit PCA with adaptive-measurement quantum tomography.