面向细粒度物体操作:基于SAM3引导的视觉运动策略,具有持久记忆学习和聚焦视觉条件化
Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning
浏览论文内容
中文总结 AI 辅助
提出SAM3引导的视觉运动框架,通过持久记忆令牌和聚焦空间-外观编码,解决细粒度物体操作中的相似物区分与干扰物鲁棒性问题,并在真实机器人任务中验证了有效性。
中文摘要 AI 辅助
细粒度物体(FO)操作要求机器人能够从视觉相似的物体中区分出指定的细粒度物体,并在场景干扰物存在的情况下可靠地执行动作。然而,场景级视觉条件化缺乏显式的物体选择,而类别级引导无法可靠地区分同一类别内的细粒度物体。我们提出了一种SAM3引导的视觉运动框架,通过持久物体记忆和聚焦视觉条件化来解决这些挑战。首先,我们引入了FO记忆驱动的SAM3(FOM-SAM3),它从有限的多视角注册图像中学习可复用的FO记忆令牌,同时保持SAM3完全冻结。通过一对多(one-vs-rest)学习,这些令牌编码了用于定位目标FO和拒绝相似替代物的持久记忆,并可存储在记忆库中。其次,我们提出了聚焦空间-外观编码(FSAE),它将FO内部局部外观特征与显式边界框坐标相结合,以条件化动作策略,包括扩散策略(DP)和基于Transformer的动作分块(ACT)。所提出的FOM-SAM3的有效性在FO-30数据集上得到了验证,该数据集包含四个粗类别中的30个物理物体。在三个真实机器人FO操作任务中,我们的FOM-SAM3引导策略展示了对干扰物的鲁棒性、对相似FO的区分能力以及对新FO的可扩展性。
英文摘要
Fine-grained object (FO) manipulation requires robots to distinguish a specified FO from visually similar objects and execute actions reliably despite scene distractors. However, scene-level visual conditioning lacks explicit object selection, while category-level guidance cannot reliably distinguish FOs within the same category. We present a SAM3-guided visuomotor policy framework that addresses these challenges through persistent object memory and focused visual conditioning. First, we introduce FO Memory-driven SAM3 (FOM-SAM3), which learns reusable FO memory tokens from limited multi-view registration images, while keeping SAM3 fully frozen. Through one-vs-rest learning, these tokens encode persistent memories for localizing target FOs and rejecting similar alternatives, which can be stored in a memory bank for reuse. Second, we propose Focused Spatial-Appearance Encoding (FSAE), which combines in-FO appearance features with explicit bounding boxes to condition action policies, including Diffusion Policy (DP) and Action Chunking with Transformers (ACT). The effectiveness of the proposed FOM-SAM3 is validated on the FO-30 dataset comprising 30 physical objects across four coarse categories. Across three real-robot FO manipulation tasks, our FOM-SAM3-guided policies demonstrated robustness against distractors, discriminative ability of similar FOs, and extendibility to new FOs.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
- State Key Laboratory of Brain Cognition and Brain-Inspired Intelligence Technology(脑认知与类脑智能技术国家重点实验室)
- Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心)
机构由 AI 辅助整理,请以论文原文为准。