对抗训练能否提升多视角VLA的泛化能力?揭示并缓解视角坍缩
Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse
浏览论文内容
中文总结 AI 辅助
本研究揭示直接对抗训练在多视角VLA中引发“视角坍缩”这一鲁棒性捷径,并引入视角交换干预以分离鲁棒感知与融合,表明AT提供选择性而非通用的分布偏移收益。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型将预训练的视觉-语言模型(VLM)适配到闭环机器人控制中,将其感知和语义能力迁移到动作预测上。尽管在分布内表现强劲,但VLA在部署偏移下往往性能下降。对抗训练(AT)提供了一种模型自适应的鲁棒性方法,无需显式预判各个偏移,但其对多视角VLA在自然分布偏移泛化上的影响尚不明确。我们使用一个直接从预训练VLM适配的多视角VLA来研究这一问题,并评估其在LIBERO-Plus七个偏移轴上的泛化性能。直接AT显著改善了仅影响第三视角的相机视点和传感器噪声这两类偏移,但对其他偏移产生混合或负面效果。受控视角干预揭示了一个我们称之为视角坍缩的惊人失败模式:直接AT可能强烈改变跨视角依赖,使得策略被腕部视角主导。这暴露了一种鲁棒性捷径:对偏移视角的表观鲁棒性可能源于减少对该视角的使用,而非对该视角更鲁棒的感知。这促使我们区分鲁棒感知(在视角内偏移下提取可靠信息)与鲁棒融合(根据可靠性调整跨视角依赖)。为减少固定的视角依赖,我们使用简单的视角交换干预并重新评估AT。在视角交换下,AT进一步改善了相机视点、传感器噪声和机器人初始状态,而其对其他偏移的效果仍为混合。我们的结果表明,多视角鲁棒性需要将感知改善与跨视角依赖变化分开考虑,且AT提供的是选择性而非通用的分布偏移收益。
英文摘要
Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial training (AT) offers a model-adaptive approach to robustness without explicitly anticipating individual shifts, but its effect on natural distribution-shift generalization in multi-view VLAs remains unclear. We study this question using a multi-view VLA directly adapted from a pretrained VLM and evaluate generalization across seven LIBERO-Plus shift axes. Direct AT substantially improves Camera Viewpoint and Sensor Noise, the two shifts affecting only the third-person view, yet produces mixed or negative effects on other shifts. Controlled view interventions reveal a surprising failure mode that we term view collapse: Direct AT can shift cross-view reliance so strongly that the policy becomes dominated by the wrist view. This exposes a \textit{robustness shortcut}: apparent robustness to a shifted view can arise from reduced use of that view rather than more robust perception of it. This motivates a distinction between robust perception, extracting reliable information under within-view shifts, and robust fusion, adapting reliance across views according to their reliability. To reduce fixed view reliance, we use a simple View Swap intervention and then re-evaluate AT. With View Swap, AT further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, while its effects remain mixed on other shifts. Our results show that multi-view robustness requires separating improved perception from changes in cross-view reliance, and that AT provides selective rather than generic distribution-shift benefits.
发表机构
- National Institute of Informatics(国立情报学研究所)
机构由 AI 辅助整理,请以论文原文为准。