AI 中文总结
该研究针对相机覆盖有限时机器人操作的视点泛化问题,提出OC-VLA++模型,通过几何引导的配对视图监督和跨视图动作等变目标提升未见视点泛化能力,性能优于原模型。
AI 中文摘要
我们提出OC-VLA++,这是OC-VLA的扩展,用于在相机覆盖范围有限的情况下实现视点泛化。OC-VLA将机器人动作建立在相机坐标系中,以对齐动作监督与视觉观测,但仅相机空间的建立仍可能过拟合到训练期间观测到的少数视点。OC-VLA++通过引入几何引导的配对视图监督和显式的跨视图动作等变目标来解决这一局限。给定来自几何相关视点的同一操作场景的配对观测,模型被训练为使其相机空间预测对应于同一机器人坐标系动作。该目标明确监督动作预测应如何跨视点变换,而非仅依赖图像级增强。实验表明,在相机覆盖范围有限的情况下,OC-VLA++对未见视点泛化有显著改进,且在相机位移增大时性能下降更平缓。这些结果确立了跨视图动作等变作为以观测为中心的动作建立的有效补充,可用于鲁棒的现实世界部署。
英文摘要
We propose OC-VLA++, an extension of OC-VLA for viewpoint generalization under limited camera coverage. While OC-VLA grounds robot actions in the camera coordinate system to align action supervision with visual observations, camera-space grounding alone can still overfit to the few viewpoints observed during training. OC-VLA++ addresses this limitation by introducing geometry-guided paired-view supervision and an explicit cross-view action-equivariance objective. Given paired observations of the same manipulation scene from geometrically related viewpoints, the model is trained such that their camera-space predictions correspond to the same robot-frame action. This objective explicitly supervises how action predictions should transform across viewpoints, rather than relying solely on image-level augmentation. Experiments demonstrate substantial improvements in unseen-view generalization under limited camera coverage, with performance degrading more gracefully under increasing camera displacement. These results establish cross-view action equivariance as an effective complement to observation-centric action grounding for robust real-world deployment.