arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06062cs.LG

付费学习,分享获益:激励式联邦多玩家老虎机

Pay to Learn, Share to Earn: Incentivized Federated Multi-Player Bandits

发表机构印度科学理工学院
查看机构详情
  • Indian Institute of Science(印度科学理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Pavamana K J, Chandramani Singh

首次发表
浏览论文内容

中文总结 AI 辅助

针对自利玩家不愿共享信息的联邦多玩家老虎机问题,提出激励感知框架及Buying-UCB算法,平衡探索与协作,理论推导遗憾与成本上界,实验验证有效性。

中文摘要 AI 辅助

联邦多玩家多臂老虎机问题建模了协作式序贯决策,其中多个玩家与一个共同的老虎机环境交互,并通过中央服务器共享信息以加速学习。现有的联邦老虎机框架通常假设所有玩家都愿意与服务器共享其本地观测。然而,这一假设在玩家为自利且在没有明确激励的情况下可能不参与协作的实际场景中往往不现实。为解决这一挑战,我们提出了一种激励感知的联邦老虎机框架,其中玩家因与服务器共享信息而获得奖励,并在从服务器购买信息时产生成本。我们开发了一种基于UCB的算法,称为Buying-UCB,该算法通过将共享激励和信息获取成本纳入学习过程,平衡了个体探索与协作学习。我们从理论上分析了所提算法,并推导出群体遗憾和购买成本的上界。我们的分析进一步刻画了完全协作联邦学习与完全独立学习之间的权衡。大量数值实验验证了理论发现,并证明了所提框架在不同协作和定价机制下的有效性。

英文摘要

Federated multi-player multi-armed bandit problems model collaborative sequential decision-making where multiple players interact with a common bandit environment and share information through a central server to accelerate learning. Existing federated bandit frameworks typically assume that all players willingly share their local observations with the server. However, this assumption is often unrealistic in practical settings where players are self-interested and may not participate in collaboration without explicit incentives. To address this challenge, we propose an incentive-aware federated bandit framework in which players receive rewards for sharing information with the server and incur costs when buying information from the server. We develop a UCB-based algorithm, termed Buying-UCB, that balances individual exploration and collaborative learning by incorporating both sharing incentives and information acquisition costs into the learning process. We theoretically analyze the proposed algorithm and derive upper bounds on the group regret and buying cost. Our analysis further characterizes the trade-off between fully collaborative federated learning and completely independent learning. Extensive numerical experiments validate the theoretical findings and demonstrate the effectiveness of the proposed framework under different collaboration and pricing regimes.

↑