发表机构
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示目标条件策略学习中目标重标记地平线增加导致的信息性诅咒,并发现蒸馏短地平线策略的输入雅可比矩阵可提升长地平线策略性能。
AI 中文摘要
学习达到目标的策略的困难常被归因于“地平线诅咒”,它表现为时间差分备份中的偏差累积和噪声优势估计。在这项工作中,我们识别了目标条件策略学习中的另一种信息性地平线诅咒,即增加目标重标记地平线会显著降低策略的泛化能力和性能。通过一系列使用预言机规划器的受控实验,我们将训练期间采样的目标地平线与测试时策略被要求达到的目标地平线解耦。即使仅在评估一系列邻近子目标时,目标条件行为克隆(BC)策略也会遭受严重的、依赖于训练地平线的性能退化,而强化学习(RL)目标可以缓解这种退化。我们将这一现象解释为动作与事后重标记目标之间的条件互信息随地平线增加而减少,并经验性地发现,在更长地平线目标上训练的BC和RL策略,其敏感性从目标信息转向状态信息,这通过策略的输入雅可比矩阵来衡量。受此观察启发,我们发现将短地平线策略的输入雅可比矩阵蒸馏到长地平线策略中,能带来显著的性能提升,尤其是在组合操作任务中。综合来看,我们的结果强调了目标重标记地平线是从离线数据学习通用策略时的一个重要考虑因素。
英文摘要
The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional informational curse of horizon in goal-conditioned policy learning, where increasing the goal relabeling horizon can significantly reduce policy generalization and performance. Through a series of controlled experiments with oracle planners, we decouple the goal horizons sampled during training from those that the policy is asked to reach at test time. Even when evaluated only on a sequence of nearby subgoals, goal-conditioned behavioral cloning (BC) policies suffer from severe, training horizon-dependent performance degradation that is mitigated by reinforcement learning (RL) objectives. We explain this phenomenon as a horizon-dependent decrease in the conditional mutual information between actions and hindsight-relabeled goals, and find empirically that both BC and RL policies trained on longer-horizon goals exhibit a shift in sensitivity from goal to state information, as measured by the policy's input Jacobians. Motivated by this observation, we find that distilling the input Jacobians of short-horizon policies into long-horizon policies yields significant performance gains, especially in combinatorial manipulation tasks. Taken together, our results highlight goal relabeling horizon as an important consideration when learning generalist policies from offline data.
Comments25 pages, 11 figures