arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有重尾奖励和信息不对称的鲁棒多智能体多臂老虎机

Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

Daphne Feng, Ricardo Parada, Lily Jiang, Sophia Yi, William Chang

arXiv 2608.10529首次发表:更新:

AI 中文总结

该研究针对重尾奖励和信息不对称的多智能体多臂老虎机问题,为三种信息不对称场景开发鲁棒分散式算法,通过实验验证了理论结果并阐明相关权衡。

AI 中文摘要

多臂老虎机问题是序列决策中的核心框架,广泛在次高斯奖励假设下被研究。但现实应用常涉及重尾奖励分布,以及分散式、信息不对称的交互。我们研究重尾奖励下的多智能体多臂老虎机,涵盖三种信息不对称场景:共同奖励下的未观测动作、独立奖励下的已观测动作、独立奖励下的未观测动作。我们为每种场景开发鲁棒分散式算法,推导的悔界几乎匹配集中式重尾速率。在帕累托分布奖励环境上的实验验证了理论发现,并阐明了三种场景中同步、协调与探索间的权衡。

英文摘要

The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. However, real-world applications often involve heavy-tailed reward distributions and decentralized, information-asymmetric interactions. We study multi-agent multi-armed bandits with heavy-tailed rewards under three information-asymmetry regimes: unobserved actions with common rewards, observed actions with independent rewards, and unobserved actions with independent rewards. We develop robust decentralized algorithms for each setting and derive regret guarantees that nearly match centralized heavy-tailed rates. Experiments on a Pareto-distributed reward environment validate our theoretical findings and illustrate the trade-offs between synchronization, coordination, and exploration across the three regimes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑