发表机构
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多智能体多臂老虎机的Top-K联合臂识别问题,设计无通信的UCB-Intervals算法,量化分析了去中心化移除通信的统计代价,证明完全非对称下样本复杂度有4倍固定惩罚。
AI 中文摘要
协作团队常需就最优的少数选项达成一致而非仅积累奖励,且需在每个成员仅能看到团队集体经验的片段时完成该任务。我们将此视为多智能体多臂老虎机中的Top-K联合臂识别问题:每一轮有M个智能体同时选择构成联合臂的个体动作,团队最终需返回平均奖励最高的K个联合臂。难点在于智能体可能无法观测其他智能体的动作、奖励或两者皆不可观测。我们处理三种可观测性场景——(A)共享奖励但动作隐藏、(B)动作可观测但奖励私有、(C)完全非对称——并设计无通信的消除算法UCB-Intervals,该算法可从各场景剩余的任意信号中重构隐式协调:(A)中的共享臂排序、(B)中的可观测偏差、(C)中的扩大置信半径。我们在固定预算和固定置信度目标中给出匹配分析,随后将三种场景整合为一个由多重性c和共识因子ρ索引的元保证。我们的核心结果是定量的而非仅算法层面的:测度变换下界表明,共享奖励识别在一个通用对数因子内是最优的,且移除通信的全部统计代价是样本复杂度中一个乘法项ρ²,在完全非对称下为固定的4倍惩罚。所得停止时间缩放为O(∑_a log(A^M/δ)/Δ_a²),固定预算误差为exp(-Θ(T/H₁)),且联合动作数A^M的依赖关系被证明是不可避免的。
英文摘要
Cooperative teams often need to agree on the best few options rather than simply accumulate reward, and they must do so while each member sees only a fragment of the team's collective experience. We study this as top-$K$ joint-arm identification in multi-agent multi-armed bandits: at every round $M$ agents simultaneously choose individual actions that compose a joint arm, and the team must ultimately return the $K$ joint arms of highest mean reward. The difficulty is that an agent may not observe the actions of others, their rewards, or either. We treat three observability regimes---(A) shared rewards with hidden actions, (B) observed actions with private rewards, and (C) full asymmetry---and design communication-free elimination algorithms (UCB-Intervals) that reconstruct implicit coordination from whatever signal each regime leaves intact: a shared arm ordering in (A), observable deviations in (B), and enlarged confidence radii under (C). We give matching analyses in both the fixed-budget and fixed-confidence objectives, then fold all three regimes into a single meta-guarantee indexed by a multiplicity $c$ and a consensus factor $ρ$. Our central result is quantitative rather than merely algorithmic: change-of-measure lower bounds show that shared-reward identification is optimal up to one universal logarithmic factor, and that the entire statistical price of removing communication is a multiplicative $ρ^2$ in sample complexity---a fixed $4\times$ penalty under full asymmetry. The resulting stopping time scales as $O\!\left(\sum_{\mathbf{a}} \frac{\log(A^M/δ)}{Δ_{\mathbf{a}}^2}\right)$ and the fixed-budget error as $\exp(-Θ(T/H_1))$, with the dependence on the joint-action count $A^M$ shown to be unavoidable.