arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于接触丰富的机器人操作的表示对齐触觉基础

Representation-Aligned Tactile Grounding for Contact-Rich Robotic Manipulation

Ruilin Chen, Jingkai Jia, Tong Yang, Xinyu Zhou, Qiao Sun, Jiangwei Zhong, Shizeng Zhang, Nuo Chen, Bailin He, Wei Li, Wenqiang Zhang

arXiv 2607.14609首次发表:更新:

发表机构

Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University; College of Intelligent Robotics and Advanced Manufacturing, Fudan University; Lenovo CTO Organization; Nanyang Technological University; TeleAl, China Telecom(复旦大学计算机科学与人工智能学院智能信息处理上海市重点实验室; 复旦大学智能机器人与先进制造学院; 联想首席技术官组织; 南洋理工大学; 中国电信天翼人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究接触丰富的机器人操作中,针对VLA策略不同表示作用下触觉监督应用位置不明的问题,通过线性探针分析找到可预测未来触觉状态的中间动作专家特征,引入LTP,实验证明表示对齐的触觉基础效果更佳。

AI 中文摘要

已引入触觉增强的视觉-语言-动作(VLA)策略用于接触丰富的操作,其中关键交互状态常对视觉隐藏。未来触觉预测是利用触觉的一种有前景的方式,它将触觉结果转化为对动作引起的接触动力学的监督。然而VLA策略包含从感知编码到运动预测等不同作用的表示,不清楚监督应应用于何处。我们将此作为表示对齐问题研究。通过线性探针分析发现,未来触觉状态从中间动作专家特征比从视觉-语言特征或最终动作状态更可预测。基于此,引入轻量级潜在触觉预测器(LTP),从已识别的中间表示预测紧凑的未来触觉嵌入。实验表明表示对齐的触觉基础优于未对齐或多接口触觉预测,凸显了触觉监督应用位置的重要性。

英文摘要

Tactile-enhanced vision-language-action (VLA) policies have been introduced for contact-rich manipulation, where critical interaction states are often hidden from vision. Future tactile prediction is a promising way to use touch because it turns tactile outcomes into supervision for action-induced contact dynamics. Yet VLA policies contain representations with different roles, from perceptual encoding to motor prediction, making it unclear where this supervision should be applied. We study this as a representation-alignment problem. Through a linear probe analysis, we find that future tactile states are most predictable from intermediate action-expert features, rather than from vision-language features or final action states. Motivated by this observation, we introduce a lightweight Latent Tactile Predictor (LTP), which predicts compact future tactile embeddings from the identified intermediate representation. By avoiding direct prediction of noisy raw tactile signals, LTP provides an action-outcome grounding signal that aligns intermediate action representations with future contact consequences. Experiments on real-world contact-rich manipulation tasks show that representation-aligned tactile grounding outperforms less aligned or multi-interface tactile prediction, highlighting the importance of where tactile supervision is applied.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑