AI 中文总结
该研究针对重尾奖励和信息不对称的多智能体多臂老虎机问题,为三种信息不对称场景开发鲁棒分散式算法,通过实验验证了理论结果并阐明相关权衡。
AI 中文摘要
多臂老虎机问题是序列决策中的核心框架,广泛在次高斯奖励假设下被研究。但现实应用常涉及重尾奖励分布,以及分散式、信息不对称的交互。我们研究重尾奖励下的多智能体多臂老虎机,涵盖三种信息不对称场景:共同奖励下的未观测动作、独立奖励下的已观测动作、独立奖励下的未观测动作。我们为每种场景开发鲁棒分散式算法,推导的悔界几乎匹配集中式重尾速率。在帕累托分布奖励环境上的实验验证了理论发现,并阐明了三种场景中同步、协调与探索间的权衡。
英文摘要
The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. However, real-world applications often involve heavy-tailed reward distributions and decentralized, information-asymmetric interactions. We study multi-agent multi-armed bandits with heavy-tailed rewards under three information-asymmetry regimes: unobserved actions with common rewards, observed actions with independent rewards, and unobserved actions with independent rewards. We develop robust decentralized algorithms for each setting and derive regret guarantees that nearly match centralized heavy-tailed rates. Experiments on a Pareto-distributed reward environment validate our theoretical findings and illustrate the trade-offs between synchronization, coordination, and exploration across the three regimes.