arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10857cs.AIcs.RO

面向地平线不变行为克隆的自监督关键帧发现

Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning

Prabin Kumar Rath, Omkar Patil, Nakul Gopalan

首次发表
浏览论文内容

中文总结 AI 辅助

针对非马尔可夫环境下行为克隆的长时程依赖问题,提出Keyframe Mnemonics自监督关键帧发现方法,在合成域和机器人操纵基准上均实现了显著性能提升与长时程泛化能力。

中文摘要 AI 辅助

非马尔可夫环境下的行为克隆(BC)是一个具有挑战性的问题,因为策略必须对长时程的上下文信息进行推理。现有策略架构依赖循环或基于注意力的机制来捕捉长期依赖关系,但循环模型存在隐藏状态崩溃和通过时间反向传播时梯度不稳定的问题,而基于注意力的模型则受限于上下文长度。为解决这些问题,我们提出Keyframe Mnemonics,一种新颖的自监督方法,通过从随机采样的过去观测中学习目标,并将其作为关键帧选择的奖励,从而发现一组信息关键的观测(即“记忆符”)。随后,我们训练一个BC策略,该策略以发现的关键帧为条件来建模动作分布。在特定任务结构假设下,我们的公式在无限时程上提供上下文保留保证,同时在策略的工作记忆中保留少量与决策相关的关键帧。我们在合成记忆域上评估了我们的方法,其中以记忆符为条件的BC策略达到100%的成功率(SR),并能泛化到远超训练的时程且性能无下降。此外,我们在记忆密集型机器人操纵基准上进行了评估,在23个任务上比最强基线实现了13.9%的平均绝对SR提升,并且在真实机器人上,在20倍更长的时程下仍保留80%的SR。代码和视频可在this https URL获取。

英文摘要

Behavior cloning (BC) in non-Markovian environments is a challenging problem because policies have to reason over contextual information over long horizons. Existing policy architectures rely on recurrent or attention-based mechanisms to capture long-term dependencies. However, recurrent models suffer from hidden-state collapse and gradient instability under backpropagation through time, while attention-based models are fundamentally limited by context length. To address these issues, we propose Keyframe Mnemonics, a novel self-supervised method that $\textit{discovers}$ a set of information-critical observations ($\textit{mnemonics}$) by learning an objective from randomly sampled past observations and using it as a reward for keyframe selection. We then train a BC policy that conditions on the discovered keyframes to model the action distribution. Under certain task-structure assumptions, our formulation provides context retention guarantees over an infinite horizon, while maintaining a small set of decision-relevant keyframes in the policy's working memory. We evaluate our method on synthetic memory domains, where mnemonic-conditioned BC policies achieve $100$% success rates (SR) and generalize to horizons orders of magnitude beyond training without performance degradation. Additionally, we evaluate on memory-intensive robot manipulation benchmark, achieving a $13.9$% average absolute SR improvement over the strongest baseline across $23$ tasks and retaining $80$% SR at $20\times$ longer horizons on a real robot. Code and videos are available at https://keyframe-mnemonics.github.io.

发表机构

  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑