arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10128cs.LGstat.ML

m-集合对抗性赌博机与赢家反馈

m-Set Adversarial Bandits with Winner Feedback

Nicolò Cesa-Bianchi, Matteo Papini

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对m-集合对抗性赌博机,在不同效用和反馈模型下推导出遗憾的上下界,揭示了设置变化对学习速率的影响,并通过实验验证了理论分析。

中文摘要 AI 辅助

我们针对不同效用(赢家奖励或奖励总和)和反馈模型(赢家索引、赢家奖励、奖励总和及其组合)下的$m$-集合对抗性赌博机,给出了遗憾的上界和下界。通过与组合赌博机和MNL赌博机的标准界比较,我们的结果揭示了设置中的细微变化如何对学习速率产生显著影响。我们的主要技术贡献是遗憾的信息论下界。在合成数据上的实验证实了我们的理论分析。

英文摘要

We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations). By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates. Our main technical contributions are the information-theoretic lower bounds on the regret. Experiments on synthetic data confirm our theoretical analyses.

发表机构

  • Università degli Studi di Milano(米兰大学)
  • Politecnico di Milano(米兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑