发表机构
Zhejiang University; Peking University; Tsinghua University; Shanghai Jiao Tong University; Nankai University; Sun Yat-sen University(浙江大学; 北京大学; 清华大学; 上海交通大学; 南开大学; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出度量交互框架,通过交互中心令牌和度量动作交互场显式建模物体与场景的度量关系,以少量参数提升多种VLA/WAM基线的操作成功率。
AI 中文摘要
视觉-语言-动作模型和世界-动作模型推进了语言条件下的机器人操作,但往往将动作、物体和场景几何之间的度量关系隐含化。人类操作将任务相关物体的语义理解与引导手部相对于物体及其周围环境运动的空间反馈相结合。受此启发,我们引入了一个度量交互框架,在共享度量尺度下,于物理笛卡尔空间中建模物体级和场景级交互。在物体层面,交互中心令牌(ICTs)显式表示相对于被操作物体的末端执行器姿态轨迹,并与动作联合去噪,提供物理基础的交互监督。在场景层面,度量动作交互场(MAIF)使用动作和ICT查询来关注度量场景点云特征,并学习几何条件下的动作修正。通过两阶段适应,我们的框架以少量额外参数和训练步骤改进了多种VLA和WAM基线。实验表明,在LIBERO和RoboTwin~2.0上平均成功率分别提升了0.80和3.59个百分点,同时在真实世界任务上提升了6.80个百分点,在其分布外变体上提升了7.45个百分点。
英文摘要
Vision-Language-Action models and World-Action Models have advanced language-conditioned robotic manipulation, yet often leave metric relations among actions, objects, and scene geometry implicit. Human manipulation combines semantic understanding of task-relevant objects with spatial feedback that guides hand motion relative to objects and their surroundings. Inspired by this, we introduce a metric interaction framework that models object-level and scene-level interactions in physical Cartesian space at a shared metric scale. At the object level, Interaction-Centric Tokens (ICTs) explicitly represent end-effector pose trajectories relative to manipulated objects and are jointly denoised with actions, providing physically grounded interaction supervision. At the scene level, the Metric Action Interaction Field (MAIF) uses action and ICT queries to attend to metric scene point-cloud features and learns geometry-conditioned action corrections. Through two-stage adaptation, our framework improves diverse VLA and WAM baselines with a small number of additional parameters and training steps. Experiments demonstrate average success-rate gains of 0.80 and 3.59 percentage points on LIBERO and RoboTwin 2.0, respectively, alongside gains of 6.80 percentage points on real-world tasks and 7.45 percentage points on their out-of-distribution variants.