GS-VLA:基于高斯溅射的冻结VLA策略的即插即用视角规范化方法
GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting
浏览论文内容
中文总结 AI 辅助
GS-VLA是首个基于3D高斯的VLA视角规范化即插即用框架,通过轻量4M参数的3D高斯规范化器提升VLA策略视角偏移鲁棒性,无需重训即可在多维度恢复性能。
中文摘要 AI 辅助
本文提出了一种轻量的即插即用框架,可提升视觉-语言-动作(Vision-Language-Action, VLA)策略对视角偏移的鲁棒性,无需对策略进行重新训练。据我们所知,这是首个直接利用基于3D高斯的新视角合成来实现VLA策略观测空间适配的方法。当前VLA的性能依赖于训练与部署阶段相机配置完全相同的隐含假设。我们的实验表明,即便是相机支架发生微小位移,在LIBERO基准测试中,最坏情况下的成功率也会从约90%降至约10%。现有方法如大规模微调或生成式数据增强,计算成本高昂且存在灾难性遗忘的风险。为解决该问题,视角偏移被重新表述为局部新视角合成问题。在局部性假设下,即相机扰动相对于工作空间保持在较小的有界区域内,视角规范化可简化为与场景和策略无关的去遮挡任务。本研究通过在冻结的VLA策略前添加一个参数规模为4M的3D高斯规范化器来实现该思路。在不修改策略权重的情况下,GS-VLA可在三个正交维度上提升性能:(1)策略架构;(2)未见任务套件;(3)扰动规模。这些结果表明,轻量视觉模块可在无需重新训练策略的情况下,恢复视角偏移导致的大部分性能损失。
英文摘要
This paper proposes a lightweight, plug-and-play framework that improves robustness to viewpoint shifts in Vision-Language-Action (VLA) policies without policy retraining. To our knowledge, this is the first approach to directly leverage 3D Gaussian-based novel-view synthesis for observation-space adaptation in VLA policies. Current VLA performance relies on the implicit assumption that training and deployment camera configurations are identical. Our experiments show that even a small displacement of the camera mount can reduce the success rate on the LIBERO benchmark from about 90% to about 10% in the worst case. Prior approaches, such as large-scale fine-tuning or generative data augmentation, are computationally expensive and risk catastrophic forgetting. To address this, viewpoint shifts are reformulated as a localized novel-view synthesis problem. Under a Locality assumption, that camera perturbations remain within a small bounded region relative to the workspace, viewpoint normalization reduces to a scene- and policy-independent disocclusion task. Our work implements this idea with a 4M-parameter 3D-Gaussian canonicalizer prepended to a frozen VLA policy. Without modifying policy weights, GS-VLA improves performance across three orthogonal axes: (1) Policy architectures, (2) Unseen task suites, and (3) Perturbation scales. These results show that a lightweight visual module can recover a large fraction of the performance lost under viewpoint shift, without policy retraining.
发表机构
- Dankook University(檀国大学)
机构由 AI 辅助整理,请以论文原文为准。