arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32048cs.LGcs.AI

交互式分布鲁棒多智能体学习与通用函数逼近

Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation

Debamita Ghosh, George K. Atia, Yue Wang

AI总结:

针对多智能体强化学习中的模型误设问题,提出RoMEX-φ框架,结合均衡探索与对偶拟合学习,利用鲁棒MADC刻画复杂度,实现次线性鲁棒遗憾,并在实验中展现更强的鲁棒性。

AI中文摘要:

模型误设是多智能体强化学习中的一个基本挑战,其中转移不确定性可能被智能体之间的策略性交互放大。分布鲁棒马尔可夫博弈(DRMGs)为解决此类不确定性提供了一个原则性框架,然而现有方法往往依赖限制性假设或难以扩展到大规模状态和联合动作空间。我们研究具有通用函数逼近和φ-散度不确定性集的一般和DRMGs中的在线学习。我们提出RoMEX-φ,一个无模型框架,将基于均衡的探索与对偶拟合学习相结合。通过鲁棒多智能体贝尔曼算子的函数对偶表示,RoMEX-φ利用中心化经验鲁棒差异,从名义交互数据中实现可处理的 worst-case 值估计。我们引入鲁棒多智能体解耦系数(robust MADC)来刻画由策略性交互和对抗性转移不确定性引起的内在探索复杂度。我们建立了由robust MADC而非显式状态和联合动作空间大小控制的次线性鲁棒遗憾保证,用内在函数类复杂度取代表格依赖性。在总变差不确定性下的可扩展一般和DRMG上的数值实验表明,RoMEX-φ对转移偏移的鲁棒性显著优于其非鲁棒对应方法,同时与精确表格鲁棒基线保持竞争力。我们的结果为具有通用函数逼近的分布鲁棒多智能体强化学习提供了一个可扩展框架。

英文摘要:

Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for addressing such uncertainty, yet existing methods often rely on restrictive assumptions or scale poorly to large state and joint action spaces. We study online learning in general-sum DRMGs with general function approximation and $ϕ$-divergence uncertainty sets. We propose RoMEX-$ϕ$, a model-free framework that integrates equilibrium-based exploration with dual fitted learning. Through a functional dual representation of the robust multi-agent Bellman operator, RoMEX-$ϕ$ enables tractable worst-case value estimation from nominal interaction data using a centered empirical robust discrepancy. We introduce the robust Multi-Agent Decoupling Coefficient (robust MADC) to characterize the intrinsic exploration complexity arising from strategic interactions and adversarial transition uncertainty. We establish sublinear robust regret guarantees governed by the robust MADC rather than explicitly by the state and joint action space sizes, replacing tabular dependence with intrinsic function-class complexity. Numerical experiments on a scalable general-sum DRMG under total variation uncertainty show that RoMEX-$ϕ$ is substantially more resilient to transition shifts than its non-robust counterpart while remaining competitive with an exact tabular robust baseline. Our results provide a scalable framework for distributionally robust multi-agent reinforcement learning with general function approximation.

补充信息

↑