arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有延迟信息共享的分散式部分可观测团队决策方法

A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing

Xiaoxing Ren, Thomas Parisini, Andreas A. Malikopoulos

arXiv 2609.26783首次发表:更新:

发表机构

Cornell University; Imperial College London; Aalborg University; University of Trieste(康奈尔大学; 帝国理工学院; 奥尔堡大学; 的里雅斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低秩潜在动力学且模型未知的分散式部分可观测团队决策问题,提出结合团队等价性与低秩表示的分散学习规划算法,各成员基于局部信息和延迟共享信息学习近似MDP并采用最小二乘值迭代,无需集中协调或训练,可逼近集中式团队最优解,并给出有限样本保证与样本复杂度界。

AI 中文摘要

我们研究了具有低秩潜在动力学和未知系统模型的分散式部分可观测团队决策问题。所提出的框架将团队理论等价性与低秩模型表示相结合,以解决在无先验转移模型知识的部分可观测马尔可夫决策过程中的合作决策问题。每个团队成员基于局部私有信息和跨团队共享的延迟公共信息做出决策。仅利用这些可用信息,每个成员学习一个近似的低秩马尔可夫决策过程,并应用最小二乘值迭代来计算其策略。这产生了一种完全分散的学习和规划算法,既不需要集中式协调器,也不需要集中式训练。我们证明了所得的成员侧解逼近集中式团队解:尽管存在部分可观测性、未知动力学和延迟公共信息,每个成员仍能恢复近似团队最优策略的相应分量。我们进一步建立了有限样本性能保证,并推导了所提出算法的相应样本复杂度界。

英文摘要

We study decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models. The proposed framework combines team-theoretic equivalence with low-rank model representations to address cooperative decision-making in partially observable Markov decision processes without prior knowledge of the transition model. Each team member makes decisions based on local private information and delayed common information shared across the team. Using only this available information, each member learns an approximate low-rank Markov decision process and applies least-squares value iteration to compute its policy. This yields a fully decentralized learning and planning algorithm that requires neither a centralized coordinator nor centralized training. We show that the resulting member-side solutions approximate the centralized team solution: despite partial observability, unknown dynamics, and delayed common information, each member recovers the corresponding component of an approximate team-optimal policy. We further establish finite-sample performance guarantees and derive a corresponding sample-complexity bound for the proposed algorithm.

Comments15 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑