发表机构
University of Liège; KTH Royal Institute of Technology; McGill University; Mila Québec AI Institute; Belerion(列日大学; 瑞典皇家理工学院; 麦吉尔大学; 魁北克人工智能研究所; Belerion)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出记忆状态评论家,利用策略自身记忆替代历史状态,实现无偏梯度并简化非对称演员-评论家结构,在视觉追逃任务中表现更优且收敛更快。
AI 中文摘要
在部分可观测马尔可夫决策过程中,最优策略通常依赖于观测历史和过去动作。当训练期间可获得额外信息(如环境的真实状态)时,非对称演员-评论家方法已成为学习此类策略的流行方法。执行时不需要的评论家被赋予访问状态的权限。仅基于状态进行条件设定的评论家通常定义不明确,并会产生有偏的策略梯度。通过同时基于状态和历史进行条件设定,历史状态评论家恢复了正确定义和无偏梯度。在本文中,我们表明,基于状态和策略自身记忆(即策略选择动作所依据的历史的内部表示)进行条件设定的评论家已经是定义明确的,并给出无偏的策略梯度,从而消除了对第二个循环近似器来估计历史的需求。我们将其称为记忆状态评论家。由此可知,基于策略记忆的评论家无需将其损失反向传播到该记忆中,即使该记忆是历史的有损编码。我们在两个四旋翼飞行器之间、跨越两种竞技场类型的基于视觉的追逃环境中评估了记忆状态评论家。追逐者是学习智能体,而逃避者每回合从固定的启发式行为池中采样。结果表明,记忆状态评论家优于历史状态评论家,且收敛更快。此外,与仅基于状态的评论家相比,它不仅是无偏的,而且在墙壁竞技场中保持轻微优势,在该环境中,演员的历史携带了特权状态单独无法提供的信息。
英文摘要
In partially observable Markov decision processes, the optimal policy generally depends on the history of observations and past actions. Asymmetric actor-critic methods have become popular to learn such policies when additional information, such as the true state of the environment, is available during training. The critic, which is not needed at execution, is given access to the state. A critic conditioned on the state alone is generally ill-defined and yields biased policy gradients. Conditioning on the state and the history, the history-state critic restores both. In this paper, we show that conditioning the critic on the state and the policy's own memory, i.e., the internal representation of the history through which the policy selects its actions, is already well-defined and gives unbiased policy gradients, removing the need for a second recurrent approximator of the history. We call it the memory state critic. It follows that a critic based on the policy's memory need not backpropagate its loss into that memory, even though the memory is a lossy encoding of the history. We evaluate the memory-state critic in a vision-based pursuit-evasion environment between two quadrotors across two arena types. The pursuer is the learning agent, and the evader is sampled per episode from a fixed pool of heuristic behaviours. The results show that the memory-state critic outperforms the history-state critic and converges faster. In addition to being unbiased compared to the state-only critic, it maintains a slight edge in the wall arena, where the actor's history carries information that the privileged state alone does not.
CommentsAccepted at the 19th European Workshop on Reinforcement Learning (EWRL 2026), Lille, France. Non-archival workshop