发表机构
Elmore Family School of Electrical and Computer Engineering; Purdue University(埃尔莫尔家族电气与计算机工程学院; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对耦合动力学环境的控制问题,提出联合马尔可夫决策过程(JMDP)的最优控制方法,定义分布型贝尔曼最优算子并证明收敛性,推导神经近似采样目标,拓展了JMDP的应用范围。
AI 中文摘要
耦合动力学环境会展示在外生随机性的共同实现下,若干可能反事实动作后续的单步结果。普通马尔可夫决策过程(MDP)形式主义可用于推理每个动作的边际分布规律,但会丢弃这些反事实结果间的依赖关系。联合马尔可夫决策过程(JMDP)形式主义则保留了这种依赖关系。已有研究确立了JMDP的形式主义并解决了其中的固定策略联合矩评估问题,本文则开发了最优控制方法:我们为JMDP定义了非参数分布型贝尔曼最优算子,并证明当诱导的边际MDP具有唯一最优策略时,其迭代会在Wasserstein距离下收敛至最优联合回报分布;对于前两阶矩,我们在更弱的条件下建立了收敛性,该条件允许多个均值最优动作存在,只要它们的平局解决方式共享一个二阶矩不动点;我们还推导了用于神经近似的采样目标。
英文摘要
Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism allows one to reason about the marginal law of each action but discards dependence across these counterfactual outcomes. The Joint Markov decision process (JMDP) formalism preserves that dependence. Prior work established the formalism and solved the fixed-policy joint moment evaluation problem in JMDPs. This paper develops optimal-control methods. We define a nonparametric distributional Bellman optimality operator for JMDPs, and prove that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law. For the first two moments, we establish convergence under a weaker condition that permits several mean-optimal actions as long as their tie resolutions share a second-moment fixed point. We also derive sampled targets for neural approximation.
Comments14 pages, 5 figures