arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08366cs.MA

局部交互多智能体MDP的可达性认证子团队分解

Reachability-Certified Subteam Decomposition for Locally Interacting Multi-Agent MDPs

发表机构香港大学 · 加州大学圣地亚哥分校 · 卡内基梅隆大学
查看机构详情
  • University of Hong Kong(香港大学)
  • University of California, San Diego(加州大学圣地亚哥分校)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiangwu Wang, Chengwei Cao, Hongyuan Tang

首次发表
浏览论文内容

中文总结 AI 辅助

针对局部交互多智能体MDP,提出可达性认证子团队分解(RCSD),结合速度界限和奖励包络形成亲和力,实现持久划分,显著降低执行遗憾,并保证最坏情况紧界。

中文摘要 AI 辅助

持久通信限制迫使多智能体系统决定在 rollout 过程中哪些智能体可以协调。仅凭当前邻近性是不够的:分离的智能体可能稍后交互,而大的成对奖励在严重折扣之前可能无法实现。我们针对具有因子化物理动力学、有限范围有序成对奖励和几乎必然运动界的有限多智能体马尔可夫决策过程,引入了可达性认证子团队分解(RCSD)。RCSD 将成对接触时间的速度限制下界与奖励包络相结合,形成当前状态亲和力。对于任何容量有效的持久划分,切割亲和力之和限定了每个不变平稳马尔可夫状态反馈策略的奖励删除误差。由此产生的切割 MDP 的团队最优策略乘积相对于集中式最优方案的遗憾最多为该证书的两倍。两个界在最坏情况下都是紧的。在一个受控的五智能体族上,相对于均匀、仅距离和仅包络划分,RCSD-Exact 将聚合归一化执行遗憾分别降低了 56.0%、28.8% 和 25.3%。一项单独的随机二维研究在 384 次精确划分和 1,440 次受限控制器评估中未发现任何界违规。精确的四智能体证据支持 RCSD 优于均匀和仅距离分组;当前接触的原始证据是边缘性的,仅包络未解决。在平衡的 8-20 智能体层中,控制器库效用是混合的:逐点成对区间支持 RCSD 优于距离和当前接触,均匀包含零,并支持仅包络和 Value-MIP 优于 RCSD。划分构建在最多 100 个智能体的中位数下保持亚秒级;最后这一结果不包括亲和力形成或 MDP 规划。

英文摘要

Persistent communication limits force a multi-agent system to decide which agents may coordinate throughout a rollout. Current proximity alone is insufficient: separated agents may interact later, whereas a large pair reward may remain unreachable until it is heavily discounted. We introduce Reachability-Certified Subteam Decomposition (RCSD) for finite multi-agent Markov decision processes with factorized physical dynamics, finite-range ordered pair rewards, and almost-sure motion bounds. RCSD combines a speed-limit lower bound on pairwise contact time with a reward envelope to form a current-state affinity. For any capacity-valid persistent partition, the sum of cut affinities bounds the reward-deletion error of every unchanged stationary Markov state-feedback policy. A product of team-optimal policies for the resulting cut MDP incurs at most twice this certificate in regret against the centralized optimum. Both bounds are worst-case tight. On a controlled five-agent family, RCSD-Exact reduces aggregate normalized execution regret by 56.0%, 28.8%, and 25.3% relative to uniform, distance-only, and envelope-only partitions. A separate stochastic two-dimensional study finds no bound violation over 384 exact-partition and 1,440 restricted-controller evaluations. Exact four-agent evidence favors RCSD over uniform and distance-only grouping; raw evidence for current contact is borderline and envelope-only is unresolved. Across balanced 8-20-agent strata, controller-library utility is mixed: pointwise paired intervals favor RCSD over distance and current contact, include zero for uniform, and favor envelope-only and Value-MIP over RCSD. Partition construction remains subsecond in median up to 100 agents; this last result does not include affinity formation or MDP planning.

补充信息

↑