AI 中文总结
研究智能体如何学习元推理能力,引入基于反应式策略不确定性得分分配计算的强化学习方法训练元推理策略,通过实证研究表明该设计能让元推理策略学会何时采用反应式策略,何时进行决策时规划,且能随反应式策略改进转向完全反应式控制。
AI 中文摘要
长期以来,人们认识到人类有能力在快速的反应式决策和较慢的审慎规划之间切换。本文研究如何让智能体学习这种被称为元推理的能力。我们将反应式决策建模为直接将状态观察映射到动作的策略,可通过强化学习或模仿学习训练,但在训练分布之外泛化性可能较差。基于模型的决策时规划能在更广泛状态集产生良好动作,但需要额外计算时间。我们引入一种强化学习方法训练元推理策略,通过反应式策略不确定性得分分配计算。在运动规划和导航环境的实证研究表明,该设计能让元推理策略学习何时反应式策略提供足够好的动作,何时需要决策时规划。此外,随着反应式策略改进,元智能体能够转向完全反应式控制。
英文摘要
It has long been recognized that humans have the ability to switch between fast, reactive decision-making and slower, deliberative planning. In this paper, we study the question of how to learn this ability, known as meta-reasoning, in artificial agents. We model reactive decision-making as a policy that directly maps state observations to actions. Such policies can be trained with reinforcement learning (RL) or imitation learning, but may generalize poorly outside of their training distribution. Alternatively, model-based decision-time planning is more likely to produce good actions across a broader set of states but requires additional computation time, which delays acting. In this work, we introduce an RL method for training a meta-reasoning policy that allocates computation by conditioning on a reactive-policy uncertainty score. This score enables it to predict when the reactive policy is likely to perform poorly and when planning is needed. We conduct an empirical study on motion planning and navigation environments, showing that this design enables the meta-reasoning policy to learn when the reactive policy provides a good-enough action versus when decision-time planning is needed. Additionally, we show that our design enables the meta-agent to shift toward fully reactive control as the reactive policy improves.
CommentsRLC 2026