发表机构
University of Toronto; Labs(多伦多大学; 2012实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出空间嫁接模块,将冻结重建特征绑定到机器人相对度量几何,通过交叉注意力注入流匹配策略,在模拟和真实机器人上广泛提升VLA和WAM操作性能。
AI 中文摘要
预训练的机器人操作策略(如视觉-语言-动作模型(VLAs)或世界-动作模型(WAMs))将交互相关的度量几何隐含化。近期空间重建的突破能够可靠地提供必要的几何信息,但其特征仅描述局部形状,而未指明其在机器人坐标系中的位置。如何最佳地将这些特征传递给预训练策略仍未解决。我们提出空间嫁接(Spatial Grafting),一种通用、轻量的空间模块,将冻结的重建特征绑定到度量、机器人相对几何上。空间嫁接构建度量锚定的空间令牌,并通过交叉注意力将其注入流匹配动作专家,而不修改主干的感知路径,从而使主干保留其预训练的全部优势。我们对其评估比任何所比较的几何感知策略更广泛:一种嫁接架构,无需针对每个主干重新设计,应用于两个VLA和两个WAM,跨越四个模拟基准,涵盖短视域操作、视觉鲁棒性、杂乱环境及长视域移动操作,并在三个真实机器人平台(单臂和双臂配置)上测试。在双臂操作基准RoboTwin 2.0上,嫁接提升了所有VLA和WAM主干。嫁接的π0.5在干净和随机场景中分别提升11.3%和15.6%,达到94.0%和92.4%,超过最强已发表的3D条件策略WAM4D(93.8%和89.9%)。随着视域延长,差距扩大:在BEHAVIOR-1K的任务上(一个以平均任务进度评分的双臂移动操作挑战),它在六个任务中的五个上超越2025年挑战赛冠军,最高提升0.47 Q分数,并在两者共同报告的三个任务上平均超过地图条件空间策略。
英文摘要
Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted $π_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.
Comments17 pages, 4 figures, 9 tables