Scaling Internal-State Policy-Gradient Methods for POMDPs
在部分可观测马尔可夫决策过程中的内部状态策略梯度方法扩展
机构 * Research School of Information Science and Engineering, Australian Nat. University, ACT 0200, Australia(信息科学与工程研究学校,澳大利亚国立大学,ACT 0200,澳大利亚) ; Panscient Pty Ltd, Adelaide, Australia(Panscient Pty Ltd,阿德莱德,澳大利亚)
AI总结 本文提出改进的策略梯度方法,用于在无限时间 horizon 设置中学习具有记忆的策略,通过直接环境模型或模拟解决大规模 POMDPs 问题。
Journal ref Proceedings of the 19th International Conference on Machine Learning (ICML 2002), 2002, pp. 3--10