arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

带反馈的公平价格:一个紧密的极小极大特征

Price of Fairness in Bandits: A Tight Minimax Characterization

Dhruv Sarkar, Soumyadeep Dutta, Sayak Ray Chowdhury

arXiv 2607.13402首次发表:更新:

发表机构

Indian Institute of Technology Kharagpur; Mohamed bin Zayed University of Artificial Intelligence; Indian Institute of Technology Kanpur(印度理工学院卡拉格浦尔分校; 穆罕默德·本·扎耶德人工智能大学; 印度理工学院坎浦尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

带反馈问题中,标准算法或使早期参与者面临不公平损失。此前\(p \geq 0\)时结果已知,\(q = -p > 0\)的严格公平机制待解决。本文用大海捞针构造证明下界,引入\textsf{UCB - HARE}算法,其遗憾与下界匹配,实验证实该算法优于均匀探索基线。

AI 中文摘要

在带反馈问题中,标准的使遗憾最小化的算法将探索视为一种分摊成本,这可能会使早期参与者在临床试验等场景中面临不公平的事前损失。近期工作通过广义\(p\)均值评估每轮预期奖励序列,在功利主义福利(\(p = 1\))、纳什福利(\(p \to 0\))和罗尔斯公平(\(p \to -\infty\))之间进行插值。虽然已知\(p \geq 0\)时的紧密保证,但严格公平机制\(q = -p > 0\)仍未解决,因为负幂均值由最小的每轮奖励主导。对于具有非负均值的\(\sigma\) - 次高斯奖励,之前最好的算法依赖均匀早期探索,遗憾为\(O(k^{(q + 1)/2}/\sqrt{T})\),而唯一的一般下界是经典的\(\Omega(\sigma\sqrt{k/T})\)。我们通过确定严格公平的确切多项式价格来弥合这一差距。使用大海捞针构造,我们证明了一个与算法无关的下界\(\Omega(\sigma\sqrt{k^{\max(1,q)}/T})\);对于\(q > 1\),这表明惩罚\(k^{q/2}\)在信息理论上是不可避免的。然后我们引入了\textsf{UCB - HARE}(调和锚定秩探索),它用由认证正均值锚保护的逆加权调和秩调度取代均匀探索。其遗憾为\(\widetilde{O}(\sigma\sqrt{k^{\max(1,q)}/T})\),与下界在对数因子上匹配。在合成实例上进行的实验证实,\textsf{UCB - HARE}优于均匀探索基线,随着\(q\)的增加收益也增加。

英文摘要

In bandit problems, standard regret-minimizing algorithms treat exploration as an amortized cost, which can expose early participants to unfair ex-ante losses in settings such as clinical trials. Recent work addresses this by evaluating the sequence of per-round expected rewards through the generalized $p$-mean, interpolating between utilitarian welfare ($p=1$), Nash welfare ($p\to0$), and Rawlsian fairness ($p\to-\infty$). Although tight guarantees are known for $p\ge0$, the strictly fair regime $q=-p>0$ remains unresolved because negative-power means are dominated by the smallest per-round rewards. For $σ$-sub-Gaussian rewards with nonnegative means, the best prior algorithm relied on uniform early exploration and achieved regret $O(k^{(q+1)/2}/\sqrt{T})$, while the only general lower bound was the classical $Ω(σ\sqrt{k/T})$. Thus it was unclear whether the extra dependence on $k$ was intrinsic to strict fairness or an artifact of uniform exploration. We close this gap by identifying the exact polynomial price of strict fairness. Using a needle-in-haystack construction, we prove an algorithm-independent lower bound $Ω(σ\sqrt{k^{\max(1,q)}/T})$; for $q>1$, this shows that the penalty $k^{q/2}$ is information-theoretically unavoidable. We then introduce \textsf{UCB-HARE} (Harmonic Anchored Rank Exploration), which replaces uniform exploration with an inverse-weighted harmonic rank schedule protected by a certified positive-mean anchor. Its regret is $\widetilde{O}(σ\sqrt{k^{\max(1,q)}/T})$, matching the lower bound up to logarithmic factors. Experiments on synthetic instances confirm that \textsf{UCB-HARE} improves over uniform-exploration baselines, with gains increasing as $q$ grows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑