arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillMemo:用于组合式具身操纵的专家引导技能记忆框架

SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tang, Jiwen Lu

arXiv 2608.05970首次发表:更新:

发表机构

Shenzhen International Graduate School, Tsinghua University; Department of Automation, Tsinghua University; Nanyang Technological University; Beijing Normal University(清华大学深圳国际研究生院; 清华大学自动化系; 南洋理工大学; 北京师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SkillMemo是一种专家引导的技能记忆框架,通过分解轨迹为技能基元并结合动态记忆库,提升了DP和VLA模型的组合泛化能力,在基准测试中达到最优性能。

AI 中文摘要

包括扩散策略(Diffusion Policy,DP)和视觉-语言-动作(Vision-Language-Action,VLA)模型在内的具身视觉运动模型,在机器人操纵基准测试中已展现出良好性能。然而,它们的潜力从根本上受限于大规模具身轨迹数据集的匮乏,导致在分布外(OOD)场景中组合泛化能力不足,且难以捕捉可复用的技能结构。为解决这一局限,我们提出Skill-Based Memory(SkillMemo)框架,该框架将长 horizon 演示隐式分解为潜在原子技能,并将技能级特征整合至动态情景记忆库中以解决组合任务。具体而言,我们首先引入基于混合专家(Mixture-of-Experts,MoE)架构的专家引导轨迹分割模块,该模块将轨迹隐式划分为由学习到的门控系数表示的不同技能基元。我们进一步设计了技能级情景记忆架构,将紧凑的技能表示存储为可检索的键值对。推理过程中,记忆库会检索最相关的技能基元,随后将其与模型当前的门控分布融合,为优化动作预测提供鲁棒的上下文先验。在模拟基准测试和真实世界操纵任务上开展的大量实验表明,SkillMemo可持续提升DP和VLA主干网络的性能,达到了当前最优性能,且优于π₀.₅,同时对未见过的任务配置展现出强大的组合泛化能力。

英文摘要

Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, leading to insufficient compositional generalization in out-of-distribution (OOD) scenarios with limited capability to capture reusable skill structures. To address this limitation, we propose Skill-Based Memory (SkillMemo) framework that implicitly decomposes long-horizon demonstrations into latent atomic skills and integrates skill-level features into a dynamic episodic memory bank for solving compositional tasks. Specifically, we first introduce an expert-guided trajectory segmentation module built upon a Mixture-of-Experts (MoE) architecture, which implicitly partitions trajectories into distinct skill primitives represented by learned gating coefficients. We further design a skill-level episodic memory architecture that stores compact skill representations as retrievable key-value pairs. During inference, the memory bank retrieves the most relevant skill primitives which are subsequently fused with the model's current gating distribution, providing a robust contextual prior to refine action predictions. Extensive experiments on the simulation benchmark and real-world manipulation tasks demonstrate that SkillMemo consistently enhances both DP and VLA backbones, achieving state-of-the-art performance and outperforming $π_{0.5}$, while exhibiting strong compositional generalization to unseen task configurations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑