发表机构
University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出多链鲁棒平均奖励MDP的向量贝尔曼理论,通过增益-偏差系统求解状态依赖的最优鲁棒奖励,并设计近似移位Halpern算法实现有限收敛。
AI 中文摘要
鲁棒平均奖励马尔可夫决策过程为不确定性下的长期性能优化提供了基础框架,其最优长期奖励可能依赖于初始状态。这种状态依赖性要求一种向量贝尔曼理论,该理论同时考虑循环类奖励和转移不确定性。我们针对具有紧致、事后$(s,a)$-矩形模糊性的有限模型发展了这样的理论。一种增益优先、偏差次之的优化原则产生了一个耦合的向量增益-偏差系统,每个有限解都能识别最优鲁棒增益,并同时从所有初始状态提供针对历史依赖对手的平稳鞍点策略。我们进一步通过平稳增益条件和典型瞬态修正的一致有界性来刻画可解性,并给出了允许不同循环类增益的充分条件。这些证书还产生了鲁棒贝尔曼算子的渐近仿射轨迹,基于此我们设计了一种鲁棒的近似移位Halpern规划算法。在有限贝尔曼可解性下,增益估计和贝尔曼位移收敛到最优增益向量,并且每个提取的贪心控制器在有限的、实例相关的预算后是平均最优的。这些结果因此将有限贝尔曼证书与状态依赖鲁棒平均奖励的无折扣规划联系起来,提供了理论理解。
英文摘要
Robust average-reward Markov decision processes provide a fundamental framework for long-term performance optimization under uncertainty, and can have optimal long-run rewards that depend on the initial state. This state dependence requires a vector Bellman theory that accounts for both recurrent-class rewards and transition uncertainty. We develop such a theory for finite models with compact, post-action $(s,a)$-rectangular ambiguity. A gain-first, bias-second optimization principle yields a coupled vector gain-bias system, and every finite solution identifies the optimal robust gain and supplies stationary saddle strategies against history-dependent opponents, simultaneously from all initial states. We further characterize solvability through stationary gain conditions and a uniform bound on canonical transient corrections, and give sufficient conditions that permit distinct recurrent-class gains. The certificates also yield asymptotically affine trajectories of the robust Bellman operator, based on which we design a robust approximately shifted Halpern planning algorithm. Under finite Bellman solvability, the gain estimates and Bellman displacements converge to the optimal gain vector, and every extracted greedy controller is average-optimal after a finite, instance-dependent budget. These results thus connect finite Bellman certificates to undiscounted planning for state-dependent robust average rewards, providing theoretical understandings.
Commentspreprint, work in progress