arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有可达性目标的回合制随机博弈中的 PAC 学习:基于期望条件距离的分散式隐私方法

PAC Learning in Turn-Based Stochastic Games with Reachability Objectives: A Decentralized Private Approach via Expected Conditional Distance

Ali Asadi, Krishnendu Chatterjee, Pavol Kebis

arXiv 2607.14877首次发表:更新:

AI 中文总结

研究回合制随机博弈中可达性目标的 PAC 学习,放宽以往公共信息和集中式学习假设,实现基于私有信息的分散式学习,引入期望条件距离参数的博弈论推广并建立多项式样本复杂度界。

AI 中文摘要

可达性是最基本的逻辑目标,但在强化学习环境中学习却十分困难,即使对于马尔可夫决策过程,若无额外假设,可达性的 PAC 学习也无法实现,在回合制随机博弈中同样如此。以往关于回合制随机博弈中 PAC 学习的文献考虑了玩家共享的公共信息和集中式学习。本文有两方面贡献:一是放宽假设,实现基于不共享的私有信息以及玩家不共享相同学习算法的分散式学习;二是引入期望条件距离参数的博弈论推广,建立了关于状态、动作、期望条件距离参数以及容错和失败概率倒数的多项式样本复杂度界。

英文摘要

Reachability is the most fundamental logical objective, yet it is notoriously difficult to learn in reinforcement learning settings: even for Markov decision processes, PAC learning of reachability is impossible without additional assumptions. This difficulty also holds in turn-based stochastic games (TBSGs), where two adversarial players interact on a finite state space. In this work, we consider turn-based stochastic games with reachability objectives. For such settings, adversarial learning, in which players are adversarial even in the learning phase, is impossible. Therefore, the goal is to consider learning, in which both players learn the unknown model together. In this spirit, previous literature on PAC learning in TBSGs considers (a)~public information shared by both players; and (b)~centralized learning, which means that players share the same learning algorithm. In this work, our contribution is two-fold. First, we relax these strong assumptions and ensure learning: (i)~with private information not shared with the other player; and (ii)~decentralized learning where the players do not share the same learning algorithm. To the best of our knowledge, this work is the first positive result for decentralized and private information learning of TBSGs with reachability objectives. Second, we introduce a game-theoretic generalization of the Expected Conditional Distance (ECD) parameter, which measures the expected length of reaching the target set. We establish a polynomial-sample complexity bound with respect to the number of states, actions, ECD parameter, and inverses of error tolerance and failure probability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑