arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多臂老虎机的纳什社会福利:轨迹级期望与高概率遗憾

Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret

Avishek Ghosh

arXiv 2610.07737首次发表:更新:

发表机构

Indian Institute of Technology, Bombay(印度理工学院孟买分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对公平多臂老虎机,提出轨迹级与高概率纳什遗憾度量,并设计RR-NCB算法,在更强度量下达到最优遗憾界。

AI 中文摘要

我们研究在纳什社会福利(NSW)目标下的公平多臂老虎机问题,该目标通过累积奖励的几何平均值来衡量性能。现有工作将纳什遗憾定义为 $\mathrm{NR}_T = \mu^\star - (\prod_{t=1}^T \mathbb{E}\mu_{I_t})^{1/T}$,其中 $\mu_{I_t}$ 是推荐臂 $I_t$ 的平均奖励,$T$ 是时间范围。由于它将对每轮边际期望应用几何平均,忽略了各轮奖励之间的联合分布,使得NSW的公平性动机在轨迹层面未得到解决。我们提出轨迹级纳什遗憾 $\widetilde{\mathrm{NR}}_T = \mu^\star - \mathbb{E}[(\prod_{t=1}^T \mu_{I_t})^{1/T}]$,它在取期望之前对完整样本路径计算几何平均,更忠实地捕捉NSW的公平性。根据Jensen不等式,$\widetilde{\mathrm{NR}}_T \geq \mathrm{NR}_T$,使其成为严格更强的度量。我们还引入了高概率纳什遗憾 $\widehat{\mathrm{NR}}_T = \mu^\star - (\prod_t \mu_{I_t})^{1/T}$,给出了公平老虎机中首个高概率遗憾界。我们的两阶段算法——轮询纳什置信界(\texttt{RR-NCB}),结合了轮询探索与纳什置信界索引策略。我们证明 $\widetilde{\mathrm{NR}}_T \leq \widetilde{\mathcal{O}}(\sqrt{k\log T/T})$,并且以概率 $1-\delta$,$\widehat{\mathrm{NR}}_T \leq \widetilde{\mathcal{O}}(\sqrt{k\log(kT/\delta)/T})$,尽管度量更强,仍匹配最优的 $\widetilde{\mathcal{O}}(\sqrt{k/T})$ 速率。最优性通过AM-GM和标准 $k$ 臂老虎机极小极大论证的下界得到保证。模拟验证了我们的理论。

英文摘要

We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards. Existing work defines Nash regret as $\mathrm{NR}_T = μ^\star - (\prod_{t=1}^T \mathbb{E}μ_{I_t})^{1/T}$, where $μ_{I_t}$ is the mean reward of the recommended arm $I_t$ and $T$ is the horizon. Since it applies the geometric mean to per-round marginal expectations, it ignores the joint distribution of rewards across rounds, leaving the NSW fairness motivation unaddressed at the trajectory level. We propose \emph{trajectory-wise Nash regret} $\widetilde{\mathrm{NR}}_T = μ^\star - \mathbb{E}[(\prod_{t=1}^T μ_{I_t})^{1/T}]$, which computes the geometric mean over complete sample paths before taking expectations, capturing NSW fairness more faithfully. By Jensen's inequality, $\widetilde{\mathrm{NR}}_T \geq \mathrm{NR}_T$, making it a strictly stronger metric. We also introduce \emph{high probability Nash regret} $\widehat{\mathrm{NR}}_T = μ^\star - (\prod_t μ_{I_t})^{1/T}$, giving the first high probability regret bounds in fair bandits. Our two-phase algorithm, Round Robin Nash Confidence Bound (\texttt{RR-NCB}), combines round robin exploration with a Nash confidence bound index policy. We show $\widetilde{\mathrm{NR}}_T \leq \widetilde{\mathcal{O}}(\sqrt{k\log T/T})$ and, with probability $1-δ$, $\widehat{\mathrm{NR}}_T \leq \widetilde{\mathcal{O}}(\sqrt{k\log(kT/δ)/T})$, matching the optimal $\widetilde{\mathcal{O}}(\sqrt{k/T})$ rate despite the stronger metrics. Optimality follows from a lower bound via AM-GM and standard $k$-armed bandit minimax arguments. Simulations validate our theory.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑