协调多人老虎机中的未知利普希茨常数
Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits
中文总结 AI 辅助
针对连续动作空间中利普希茨常数未知的协作多智能体老虎机问题,设计适配三种信息结构的算法,证明共享奖励或可观测动作可免费达成离散化一致性,否则通过抖动量化估计值达成且不增加悔的主阶成本。
中文摘要 AI 辅助
受去中心化应用的启发,我们研究当利普希茨(Lipschitz)常数未知时,连续(利普希茨)动作空间中的协作多智能体老虎机问题。我们考虑三种信息结构:(A)动作未被观测但奖励共享;(B)动作被观测且奖励独立;(C)动作未被观测且奖励独立。针对每种情况,我们设计并分析一种算法,该算法会估计利普希茨常数,选择联合动作空间的离散化方案,并将协作老虎机方法应用于由此产生的离散问题。学习开始后玩家不再通信,因此核心难点在于他们必须从各自的数据中达成相同的离散化方案。我们证明了悔界保证:共享奖励和可观测动作可免费提供这种一致性,而在缺少这两者的情况下,仍可通过对估计值进行抖动量化来获得一致性,且不会增加悔的主阶成本。
英文摘要
Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown. We consider three information structures: (A)~unobserved actions with common rewards, (B)~observed actions with independent rewards, and (C)~unobserved actions with independent rewards. In each case we design and analyze an algorithm that estimates the Lipschitz constant, chooses a discretization of the joint action space, and applies a cooperative bandit method to the induced discrete problem. Players never communicate once learning starts, so the central difficulty is that they must reach the \emph{same} discretization from their own data. We prove regret guarantees showing that common rewards and observable actions each supply this agreement for free, and that in their absence agreement can still be bought, through a dithered quantization of the estimate, at no cost in the leading order of the regret.