状态轨迹理由作为强化学习中的辅助任务
State Trace Rationale As Auxiliary Task in Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
提出STRAT辅助任务,让强化学习智能体预测自身状态文本轨迹,结合地标、路线和概览知识,在60个稀疏奖励任务中解决标准RL失败的环境,并压缩状态表示、防止秩坍缩,同时提供可读的信念描述。
中文摘要 AI 辅助
我们提出STRAT,一种辅助任务,训练深度强化学习(RL)智能体预测自身状态的简短文本轨迹。受人类空间导航启发,该描述结合地标、路线和概览知识,跟踪智能体的位置、库存、目标和即时进度。环境规则在线生成此文本,无需人工标注。我们的方法在标准策略上增加一个辅助头。在60个稀疏奖励的XLand-MiniGrid任务中,STRAT解决了标准RL完全失败的复杂环境,同时压缩状态表示并防止秩坍缩。除了性能提升,预测轨迹在每一步都提供了智能体信念的可读描述,且无需额外成本。
英文摘要
We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.
发表机构
- University of the Witwatersrand(金山大学)
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。