发表机构
Tsinghua Shenzhen International Graduate School, Tsinghua University; Southern University of Science and Technology; School of Future Technology, Harbin Institute of Technology(清华大学深圳国际研究生院; 南方科技大学; 哈尔滨工业大学未来技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对接触丰富操作的视觉反馈不足问题,提出基于VDF的VISTA-Policy,通过整合物理感知编码引擎等模块,在多任务上优于纯视觉与触觉基准,具分布外泛化及抗干扰能力。
AI 中文摘要
接触丰富的操作需要精准的交互反馈。尽管以视觉为中心的模仿学习十分流行,但外部视觉观测仅能提供关于接触状态的间接且模糊的线索,尤其在遮挡或夹爪与物体的交互较为细微时;专用的触觉或力传感器可提供丰富的接触信息,但会引入额外的硬件复杂度、校准要求与部署成本。为弥合这一差距,我们提出VISTA-Policy,一种利用视觉变形场(Visual Deformation Field,VDF)的模仿学习范式,其中VDF是柔顺夹爪的三维位移表示,作为高维视觉-物理反馈。该框架整合了三部分:1)用于实时VDF解码的物理感知编码引擎;2)用于分离真实交互信号的能量聚合去噪机制;3)带有增量夹爪动作的变形增强策略网络,以实现精准的闭环校正。在跨尺度物体抓取、拧瓶盖及书法书写任务上的大量评估表明,VISTA-Policy的性能优于强大的纯视觉基准模型3D Diffusion Policy和触觉基准模型。VISTA-Policy还展现出对未见物体尺度的显著分布外泛化能力,以及对动态干扰的鲁棒性,为非结构化环境中通用精细操作提供了一条耐用且具成本效益的路径。项目视频与补充材料可在以下网址获取:this https URL。
英文摘要
Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provide rich contact information but introduce additional hardware complexity, calibration requirements, and deployment costs. To bridge this gap, we propose VISTA-Policy, an imitation learning paradigm that utilizes the Visual Deformation Field (VDF), a 3D displacement representation of a compliant gripper, as high-dimensional visuo-physical feedback. The framework integrates: 1) a Physics-Aware Encoding Engine for real-time VDF decoding; 2) an Energy Aggregation Denoising Mechanism to isolate true interaction signals; and 3) a Deformation-Augmented Policy Network with incremental gripper actions for precise closed-loop correction. Extensive evaluations on Cross-Scale Object Grasping, Cap Unscrewing, and Calligraphy Writing demonstrate that VISTA-Policy outperforms the strong pure-vision baseline 3D Diffusion Policy and the tactile baseline. VISTA-Policy further demonstrates substantial out-of-distribution generalization to unseen object scales and robustness against dynamic disturbances, offering a durable and cost-effective route toward general-purpose fine-grained manipulation in unstructured environments. Project videos and supplementary materials are available at: https://sites.google.com/view/vista-policy.