发表机构
Georgia Institute of Technology; University of Massachusetts Amherst(佐治亚理工学院; 马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究对抗性多目标老虎机中坐标信息对帕累托遗憾的影响,证明未知易坐标无法改善最坏情况遗憾,并开发自适应Poly-INF算法达到匹配的极小极大最优速率。
AI 中文摘要
对抗性多目标老虎机有望帮助我们优化选择(臂),其奖励是由对手选择的多维向量,其性能通过帕累托遗憾来衡量。我们将损失定义为一减去奖励,并通过臂在某个坐标上的最小累积损失来衡量该坐标的容易程度,当该值较小时称该坐标更容易。现有工作表明,理论上更易的坐标可能降低帕累托遗憾。然而,在实践中,人们可能不知道哪个坐标更容易。在消极方面,我们表明这种信息的缺失消除了这种可能性:较小的累积损失并不能改善帕累托遗憾的最坏情况阶数。具体来说,设$L_d$为$T$轮中沿坐标$d$的最小累积损失。对于$K\ge4$个臂、$T\ge6$轮以及至少2个坐标,我们证明极小极大期望帕累托遗憾为$\Omega(\min\{T-L_0,\sqrt{K(T-L_0)}\})$。它随$L_0$单调递减,即使$L_0=\min_d L_d$本身已知也是如此。在积极方面,这一结果促使了其他坐标(不仅仅是容易的坐标)可能足以达到帕累托遗憾的最优速率的可能性。当$L_0$已知时,我们将Poly-INF应用于一个固定坐标,并得到帕累托遗憾的上界,该上界表现出相同的阶数,从而与下界匹配。在没有这种知识的情况下,我们开发了Poly-INF的奖励加倍版本,该版本适应这一未知量,同时仍达到匹配的极小极大速率。另一个含义是它没有额外的$\log T$因子,并且与坐标数量无关。
英文摘要
Adversarial multi-objective bandits hold the potential to help us optimize choices (arms) whose reward is a multidimensional vector chosen by an adversary and whose performance is measured by Pareto regret. We define loss as one minus reward and measure the easiness of a coordinate by the smallest cumulative loss of the arms on it, and call the coordinate easier when this quantity is smaller. Existing work suggests that in theory an easier coordinate may reduce Pareto regret. However, in practice, one may not know which coordinate is easier. On the negative side, we show that this lack of information eliminates the possibility: a smaller cumulative loss does not improve the worst-case order of Pareto regret. Precisely, let \(L_d\) be the smallest cumulative loss along coordinate $d$ over $T$ rounds. For \(K\ge4\) arms, \(T\ge6\) rounds, and at least 2 coordinates, we prove that the minimax expected Pareto regret is \(Ω(\min\{T-L_0,\sqrt{K(T-L_0)}\})\). It is monotonically decreasing in \(L_0\), even when \(L_0=\min_d L_d\) itself is known. On the positive side, this result motivates the possibility that other coordinates, not just the easy one, may suffice to attain the optimal rate of Pareto regret. When $L_0$ is known, we apply Poly-INF to a fixed coordinate and obtain an upper bound on Pareto regret that exhibits the same order and thus matches the lower bound. Without such knowledge, we develop a reward-doubling version of Poly-INF that adapts to this unknown quantity while still attaining the matching minimax rate. Another implication is that it has no extra \(\log T\) factor and is independent of the number of coordinates.