arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16815cs.RO

重新思考视觉运动策略中的视觉具身依赖性

Rethinking Visual Embodiment Dependence in Visuomotor Policies

Hongjie Fang, Yuxuan Lu, Chenxi Wang, Haoxiang Qin, Shirun Tang, Zihao He, Shangning Xia, Jingjing Chen, Wanxi Liu, Shiquan Wang, Cewu Lu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出通过3D点云中的具身规范化,用规范末端执行器表示替代原始具身,以构建而非消除视觉具身依赖性,从而提升视觉运动策略的泛化与人到机器人迁移性能。

中文摘要 AI 辅助

视觉运动策略同时观察任务场景和动作具身,使得具身特定的视觉线索能够影响动作预测。我们将这一现象研究为视觉具身依赖性(VED),并通过跨代表性策略的线索冲突干预表明,可见的机器人配置可能成为任务进展的捷径。我们主张不应消除VED,而应围绕支持控制和泛化的具身信息来构建它。我们通过在3D点云中进行具身规范化来实现这一点,用规范末端执行器表示(CER)替换原始具身,该表示保留了控制相关的几何形状,同时抽象了具身特定的形态。其可编辑形式进一步支持针对不熟悉机器人配置的配置去相关增强。实验表明,具身规范化显著改善了无需机器人演示的人到机器人策略迁移,而仅移除具身而不保留控制相关几何形状则不足。我们进一步发现,当机器人配置与任务进展脱钩时,CER本身可能成为配置捷径;配置去相关增强缓解了这一失败模式,并在不牺牲已见配置性能的情况下恢复了稳健恢复。总之,这些结果表明,稳健的视觉运动学习受益于构建而非移除视觉具身信息。项目网站:此https URL

英文摘要

Visuomotor policies observe both the task scene and the acting embodiment, allowing embodiment-specific visual cues to influence action prediction. We study this phenomenon as visual embodiment dependence (VED) and show, through cue-conflict interventions across representative policies, that visible robot configuration can become a shortcut to task progress. Rather than eliminating VED, we argue that it should be structured around embodiment information that supports control and generalization. We realize this through embodiment canonicalization in 3D point clouds, replacing the original embodiment with a canonical end-effector representation (CER) that preserves control-relevant geometry while abstracting embodiment-specific morphology. Its editable form further enables configuration-decorrelation augmentation for unfamiliar robot configurations. Experiments show that embodiment canonicalization substantially improves human-to-robot policy transfer without robot demonstrations, while simply removing the embodiment is insufficient without preserving control-relevant geometry. We further find that CER itself can become a configuration shortcut when robot configuration becomes decoupled from task progress; configuration-decorrelation augmentation mitigates this failure mode and restores robust recovery without sacrificing performance on seen configurations. Together, these results show that robust visuomotor learning benefits from structuring, rather than removing, visual embodiment information. Project website: https://tonyfang.net/ved

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • FORTE Lab(FORTE实验室)
  • Noematrix(诺玛矩阵)
  • Flexiv(非夕科技)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑