发表机构
Shenzhen International Graduate School, Tsinghua University; Department of Automation, Tsinghua University; Nanyang Technological University; Beijing Normal University(清华大学深圳国际研究生院; 清华大学自动化系; 南洋理工大学; 北京师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SkillMemo是一种专家引导的技能记忆框架,通过分解轨迹为技能基元并结合动态记忆库,提升了DP和VLA模型的组合泛化能力,在基准测试中达到最优性能。
AI 中文摘要
包括扩散策略(Diffusion Policy,DP)和视觉-语言-动作(Vision-Language-Action,VLA)模型在内的具身视觉运动模型,在机器人操纵基准测试中已展现出良好性能。然而,它们的潜力从根本上受限于大规模具身轨迹数据集的匮乏,导致在分布外(OOD)场景中组合泛化能力不足,且难以捕捉可复用的技能结构。为解决这一局限,我们提出Skill-Based Memory(SkillMemo)框架,该框架将长 horizon 演示隐式分解为潜在原子技能,并将技能级特征整合至动态情景记忆库中以解决组合任务。具体而言,我们首先引入基于混合专家(Mixture-of-Experts,MoE)架构的专家引导轨迹分割模块,该模块将轨迹隐式划分为由学习到的门控系数表示的不同技能基元。我们进一步设计了技能级情景记忆架构,将紧凑的技能表示存储为可检索的键值对。推理过程中,记忆库会检索最相关的技能基元,随后将其与模型当前的门控分布融合,为优化动作预测提供鲁棒的上下文先验。在模拟基准测试和真实世界操纵任务上开展的大量实验表明,SkillMemo可持续提升DP和VLA主干网络的性能,达到了当前最优性能,且优于π₀.₅,同时对未见过的任务配置展现出强大的组合泛化能力。
英文摘要
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, leading to insufficient compositional generalization in out-of-distribution (OOD) scenarios with limited capability to capture reusable skill structures. To address this limitation, we propose Skill-Based Memory (SkillMemo) framework that implicitly decomposes long-horizon demonstrations into latent atomic skills and integrates skill-level features into a dynamic episodic memory bank for solving compositional tasks. Specifically, we first introduce an expert-guided trajectory segmentation module built upon a Mixture-of-Experts (MoE) architecture, which implicitly partitions trajectories into distinct skill primitives represented by learned gating coefficients. We further design a skill-level episodic memory architecture that stores compact skill representations as retrievable key-value pairs. During inference, the memory bank retrieves the most relevant skill primitives which are subsequently fused with the model's current gating distribution, providing a robust contextual prior to refine action predictions. Extensive experiments on the simulation benchmark and real-world manipulation tasks demonstrate that SkillMemo consistently enhances both DP and VLA backbones, achieving state-of-the-art performance and outperforming $π_{0.5}$, while exhibiting strong compositional generalization to unseen task configurations.