发表机构
Montanuniversität Leoben(莱奥本矿业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InfiNoVA通过将多相机演示重建为三维高斯表示并渲染密集新视角,增强机器人策略的视点不变性,在未见视角下成功率提升5.4倍。
AI 中文摘要
视觉-语言-动作(VLA)策略通常强烈依赖于训练期间看到的相机视角,当从未见过的视角部署时,会导致性能显著下降。从足够多样化的物理视角收集演示数据成本高昂,且仅能提供对视角空间的稀疏覆盖。我们引入了InfiNoVA,一种数据增强框架,它将同步的多相机演示转换为几何一致的密集训练视图分布。InfiNoVA将每条操作轨迹重建为时变的三维高斯表示,并从采样的相机姿态渲染新的观测,同时保留原始状态-动作对应关系。这种显式场景表示提高了帧级保真度和时间一致性,同时减少了在生成式新视角合成中观察到的任务关键幻觉。在四个真实世界操作任务中,使用InfiNoVA训练的策略在未见过的随机视角下,平均成功率比基于VISTA的增强和未增强策略高出5.4倍。InfiNoVA还比直接在所有五个物理相机视图上训练实现了1.7倍更高的成功率。这些结果表明,密集的、几何基础扎实的视角增强为无需修改底层策略架构即可实现相机鲁棒的机器人策略提供了一条实用途径。
英文摘要
Vision-Language-Action (VLA) policies often rely strongly on the camera viewpoints seen during training, causing substantial performance degradation when deployed from unseen perspectives. Collecting demonstrations from sufficiently diverse physical viewpoints is expensive and still provides only sparse coverage of the viewpoint space. We introduce InfiNoVA, a data-augmentation framework that converts synchronized multi-camera demonstrations into a dense distribution of geometrically consistent training views. InfiNoVA reconstructs each manipulation trajectory as a time-varying 3D Gaussian representation and renders novel observations from sampled camera poses while preserving the original state-action correspondence. This explicit scene representation improves frame-level fidelity and temporal consistency while reducing task-critical hallucinations observed in generative novel-view synthesis. Across four real-world manipulation tasks, policies trained with InfiNoVA achieve 5.4x higher average success under unseen randomized viewpoints than both VISTA-based augmentation and the unaugmented policy. InfiNoVA further achieves 1.7x higher success than training directly on all five physical camera views. These results show that dense, geometrically grounded viewpoint augmentation provides a practical route toward camera-robust robot policies without modifying the underlying policy architecture.