arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17929cs.AIcs.LO

用于鲁棒马尔可夫决策过程的自适应策略组合

Adaptive Policy Portfolios for Robust Markov Decision Processes

Kasper Engelen, Sebastian Junges, Guillermo A. Pérez, Marnix Suilen

首次发表
浏览论文内容

中文总结 AI 辅助

针对鲁棒马尔可夫决策过程的保守性问题,研究自适应策略组合,通过复杂性理论分析其验证与合成的计算难度,并提出可适配运行时专业化的离线组合构建方法。

中文摘要 AI 辅助

鲁棒马尔可夫决策过程(RMDP)针对一组合理的转移函数优化单个策略,当未知动态固定且部署后变得部分可识别时,这种方法可能过于保守。我们研究自适应策略组合:离线合成的有限组无记忆随机策略,与轻量级在线选择器配对。鲁棒悔恨是衡量组合质量的自然指标:对于每个合理环境,它衡量组合中最佳成员相对于已知该环境时最优策略的损失。Ghavamzadeh等人(2016)研究了相关的悔恨目标,重点关注安全策略改进的近似和松弛。我们对组合的验证和合成进行了复杂性理论分析:在非循环(s,a)矩形RMDP中,验证给定组合对于确定性组合已是∀ℝ-完全问题;对于一般有理多面体,即使折扣固定且动态非循环,合成一元有界大小的组合也是∃∀ℝ-完全问题;单策略情况在组合和代数上均已困难。最后,我们提出一种可适配运行时专业化的离线组合构建方法。

英文摘要

Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and become partially identifiable after deployment. We study adaptive policy portfolios: finite sets of memoryless randomized policies synthesized offline and paired with a lightweight online selector. Robust regret is a natural measure of portfolio quality: for each plausible environment, it measures the loss of the best portfolio member relative to the policy that would have been optimal had that environment been known. Related regret objectives were studied by Ghavamzadeh et al. (2016) with an emphasis on approximations and relaxations for safe policy improvement. We give a complexity-theoretic account of portfolio certification and synthesis. Certifying a given portfolio is $\forall\mathbb{R}$-complete already for deterministic portfolios in acyclic (s,a)-rectangular RMDPs. Synthesizing a portfolio of unary-bounded size is $\exists\forall\mathbb{R}$-complete for general rational polytopes, even with fixed discount and acyclic dynamics. The single-policy case is already hard, both combinatorially and algebraically. Finally, we present an offline portfolio construction that is amenable to runtime specialization.

↑