arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboMP-DINOv2:用于鲁棒机器人操作的提示而非过滤器

RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation

Han Qi, Heng Yang

arXiv 2609.25506首次发表:更新:

发表机构

Harvard School of Engineering and Applied Sciences; Harvard University(哈佛大学工程与应用科学学院; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器人操作中的视觉泛化问题,提出RoboMP-DINOv2,将掩码作为空间提示而非过滤器,结合颜色随机化,在模拟实验中显著提升鲁棒性。

AI 中文摘要

机器人操作策略必须在视觉变化下进行泛化,同时保留与动作相关的场景上下文。通用视觉编码器并非针对视觉运动控制而定制,而基于对象的方法通常使用分割掩码作为硬过滤器,丢弃可能有用的上下文。我们提出RoboMP-DINOv2(机器人掩码提示DINOv2),一种全场景视觉编码器,将掩码视为空间提示而非可见性过滤器。它从完整观测中提取密集的DINOv2特征,在掩码位置注入学习的区域特定嵌入,并联合上下文化提示和未提示的令牌以进行动作预测。我们进一步引入掩码区域颜色随机化(MCR)以提高外观鲁棒性,产生RoboMP-DINOv2-MCR。在七个模拟操作设置中,RoboMP-DINOv2在空间偏移下达到60.7%的成功率,在场景杂乱下达到59.7%,而基于DINOv2的扩散策略分别为50.7%和41.0%。在未见过的对象颜色下,RoboMP-DINOv2-MCR达到72.5%的成功率,而最强的颜色随机化基线为35.1%。额外的实验和表示分析表明,在保留行为相关场景信息的同时,鲁棒性得到提高。代码可在该https URL获取。

英文摘要

Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑