arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07228cs.LGmath.OC

部分可观测性下学习的表现比策略类更差:闭式分析

Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis

Idil Gözel

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过闭式分析发现,部分可观测强化学习中学习偏差是性能差的主因,调整演员-评论家的前瞻步数可解决问题,提供了理论依据且经实验验证。

中文摘要 AI 辅助

当强化学习智能体无法观测完整状态时,我们通常会将问题归咎于其策略:智能体无法观测足够信息以表征良好的策略。但我们证明,在可求解的场景中,更大的问题出在别处。即使存在良好的策略,且智能体的值函数具有足够的表达能力来精确描述该策略,学习过程仍会收敛到远差于该策略的结果。我们研究了一个部分可观测的线性二次问题,其中标准的演员-评论家(actor-critic)学习者可通过闭式形式求解。在默认设置下,智能体可表征的最优策略已接近最优,其成本比观测全部状态的理想控制器高10.4%,但学习过程并未找到该策略,算法最终收敛到的策略比其可获得的最优策略差35%,且我们能明确指出问题所在及原因。原因在于评论家(critic)学习过程中存在偏差,而非演员(actor)的表达能力受限。由于智能体无法将观测到的内容归因于其无法观测的状态部分,评论家会将这种无法解释的变化误读为其值估计中的尖锐曲率,进而导致演员跟随该误差偏离最优。我们推导了所得策略、其成本以及可消除该问题的设计选择(即学习者在信任自身值估计前的前瞻步数)的闭式表达式。深度强化学习实验与这些预测高度吻合,值得注意的是,为智能体提供过去观测的记忆并无帮助,而改变前瞻步数则有效。

英文摘要

When a reinforcement learning agent cannot observe the full state, we usually blame its policies: it cannot see enough to represent a good one. We show that in a solvable case the bigger problem lies elsewhere. Even when a good policy is available and the agent's value function is expressive enough to describe it exactly, learning still ends up somewhere far worse. We study a partially observed linear-quadratic problem in which a standard actor-critic learner can be solved in closed form. At our default setting the best policy the agent can represent is already close to optimal, costing 10.4% more than the ideal controller that observes everything. Learning does not find it. The algorithm instead comes to rest at a policy that is 35% worse than the best one available to it, and we can say exactly where and why. The cause is a bias in what the critic learns rather than a limit on what the actor can express. Because the agent cannot attribute what it sees to the part of the state it cannot observe, the critic misreads that unexplained variation as sharp curvature in its own value estimates, and the actor follows that error away from the optimum. We derive closed-form expressions for the resulting policy, for its cost, and for the one design choice that removes the problem, which is how far the learner looks ahead before trusting its own value estimates. Deep reinforcement learning experiments follow these predictions closely. Notably, giving the agent memory of past observations does not help, while changing how far it looks ahead does.

发表机构

  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑