arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05456stat.MLcs.LGmath.OC

网络化MDP中无模型强化学习的泰勒表示

Taylor Representations for Model-Free RL in Networked MDPs

  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Salah Chikhi, Abdelhaq Chaoui, Asuman Ozdaglar, Saurabh Amin

AI总结:

本文提出基于泰勒展开的局部评论家表示,用于网络化MDP的无模型强化学习,并设计可扩展的演员-评论家算法,在三个控制基准上匹配或超越谱方法。

AI中文摘要:

在网络化马尔可夫决策过程中,转移动态通常是未知的,且状态-动作空间随智能体数量迅速增长。在此背景下,泰勒表示自然地逼近$Q$-函数,但对$N$个智能体进行朴素的$n$阶展开需要$\Theta(N^n)$个系数。我们在平滑的期望未来局部奖励及受控导数的条件下证明了这些展开的合理性。在此条件下,有限速度的信息传播和折扣意味着局部评论家泰勒系数随到最远涉及智能体的图距离呈指数衰减。丢弃远距离智能体系数并进行边缘化,可得到可扩展的局部泰勒表示,其误差界由图局部性控制。基于这些表示,我们提出了一种可扩展的无模型演员-评论家算法,并为线性LSTD评论家建立了有限样本评论家及近平稳性保证。随后,我们引入了一种更具表达力的神经TD参数化。与先前的构造性谱方法不同,我们的方法涵盖了无法获取已知局部动态映射的设置,例如隐藏的切换线性-二次调节。在三个控制基准测试中,我们的方法在匹配或超越谱基线的同时,能高效扩展到大型图。

英文摘要:

In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate $Q$-functions, but a naive order-$n$ expansion over $N$ agents requires $Θ(N^n)$ coefficients. We justify these expansions under smooth expected future local rewards with controlled derivatives. Under this condition, finite-speed information propagation and discounting imply that local-critic Taylor coefficients decay exponentially with the graph distance to the farthest agent involved. Discarding distant-agent coefficients and marginalizing then yield scalable local Taylor representations with a bound controlled by graph locality. Building on these representations, we propose a scalable model-free actor--critic algorithm, establishing finite-sample critic and near-stationarity guarantees for a linear LSTD critic. We then introduce a more expressive neural TD parameterization. Unlike prior constructive spectral methods, our approach covers settings without access to a known local dynamics map, such as hidden switched linear--quadratic regulation. Across three control benchmarks, our method matches or outperforms spectral baselines while scaling efficiently to large graphs.

↑