arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIM-VLA:从夹爪电机反馈学习物理交互表征

MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback

Jaeyoung Lee, Jiyeon Koo, Taehwa Kim, Yerin Cha, Andrew Jaeyong Choi

arXiv 2610.08425首次发表:更新:

发表机构

Gachon University(嘉泉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MIM-VLA通过编码夹爪电机反馈为交互令牌,仅调节SmolVLA的夹爪动作通路,在真实场景中提升交互阻力选择准确率至75.0%,无需额外触觉传感器。

AI 中文摘要

视觉-语言-动作(VLA)策略主要从视觉观测和机器人状态推断抓取动作,但并未显式地表征接触后观察到的物理响应。我们提出MIM-VLA,一种基于电机反馈的架构,它将最近的夹爪电流、位置、速度和信号有效性编码为128维的交互令牌。仅使用电机的电机交互模块(MIM)使用人工审核的接触和交互阶段标签进行预训练,然后仅调节SmolVLA的夹爪动作通路;手臂动作和位置控制接口保持不变。同一令牌支持MEM选择器VLM,该VLM比较候选交互并产生基于证据的选择和解释。我们在三个真实世界场景中评估MIM-VLA:比较视觉上不同物体的交互阻力,通过主动探测消除视觉上相似的真实物体和复制品的歧义,以及轻柔抓取易碎物体(包括未见过的实例)。在13个物体对中,MIM-VLA在75.0%的试验中选择了阻力更高的物体,而SmolVLA基线为48.8%。对于评估的任务,该方法使用夹爪已有的电机反馈,不需要额外的触觉阵列、力-力矩传感器、校准的力估计或直接电流控制。

英文摘要

Vision-language-action (VLA) policies infer grasp actions primarily from visual observations and robot state, but do not explicitly represent the physical response observed after contact. We present MIM-VLA, a motor-feedback-based architecture that encodes recent gripper current, position, velocity, and signal validity as a 128-dimensional interaction token. A motor-only Motor Interaction Module (MIM) is pretrained with human-reviewed contact and interaction-phase labels and then conditions only the gripper-action pathway of SmolVLA; arm actions and the position-control interface remain unchanged. The same token supports the MEM selector VLM that compares candidate interactions and produces evidence-conditioned selections and explanations. We evaluate MIM-VLA in three real-world settings: comparing the interaction resistance of visually different objects, disambiguating visually similar real and replica objects through active probing, and gently grasping fragile objects, including held-out instances. Across 13 object pairs, MIM-VLA selects the higher-resistance object in 75.0% of trials, compared with 48.8% for the SmolVLA baseline. For the evaluated tasks, the approach uses motor feedback already available from the gripper and does not require an additional tactile array, force-torque sensor, calibrated force estimate, or direct current control.

Comments9 pages, 7 figures. Jaeyoung Lee and Jiyeon Koo contributed equally

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑