发表机构
Southeast University; National Center of Technology Innovation for EDA(东南大学; 国家EDA技术创新中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CoRe-VLA提出即插即用框架,通过重建点云并渲染训练视角观测,无需额外数据或微调,恢复相机移位下的跨视角协调,显著提升任务成功率。
AI 中文摘要
VLA将预训练的视觉-语言表示与动作生成相结合,实现跨多样化任务的语言引导控制,已成为具身智能的主流范式。然而,多项研究报道了VLA在相机移位下任务成功率显著下降,揭示了限制可靠部署的关键脆弱性。为解决此脆弱性,现有方法收集同一场景不同视角的配对观测,以微调VLA或训练视觉适应模块。不幸的是,它们需要额外的数据收集和VLA训练成本。在本文中,我们首先识别了外部相机移位下的“跨视角协调崩溃”:当失去全局视角时,机器人可能过度依赖腕部视角线索,从而以错误顺序执行子任务。受此启发,我们提出CoRe-VLA,一个即插即用框架,既不需要额外的多视角数据收集,也不需要VLA微调,可与现有VLA集成。它重建场景点云,并从VLA的训练视角渲染观测,以恢复跨视角协调。在CoRe-VLA中,渲染到相机(R2C)恢复减少了渲染引起的视觉退化,而执行轨迹条件对齐(ETCA)减少了异步执行期间的机器人空闲时间并缓解运动冲突。在5个真实机器人任务、LIBERO-100和LIBERO-Plus上的实验表明,CoRe-VLA在相机移位下显著提高了主流VLA的任务成功率。例如,在真实机器人环境中1.6米相机移位下,CoRe-VLA将PI0.5的成功率从13.3%提升至83.3%。
英文摘要
VLAs combine pretrained vision-language representations with action generation to enable language-guided control across diverse tasks, becoming a mainstream paradigm in embodied intelligence. However, multiple studies have reported VLA's substantial declines in task success under camera shifts, revealing a key vulnerability that limits reliable deployment. To address this vulnerability, existing methods collect paired observations of the same scene from different viewpoints to fine-tune the VLA or train visual adaptation modules. Unfortunately, they require additional data collection and VLA training costs. In this paper, we first identify \emph{cross-view coordination breakdown} under external camera shifts: the robot may rely too heavily on wrist-view cues and consequently execute subtasks in the wrong order when losing global view. Motivated by this, we propose CoRe-VLA, a plug-and-play framework requiring neither additional multi-view data collection nor VLA fine-tuning, which can incorporate with exsiting VLAs. It reconstructs a scene point cloud and renders the observation from the VLA's training viewpoint to restore cross-view coordination. In CoRe-VLA, Render-to-Camera (R2C) Restoration reduces rendering-induced visual degradation, while Execution-Trajectory-Conditioned Alignment (ETCA) reduces robot idle time and mitigates motion conflicts during asynchronous execution. Experiments on 5 real-robot tasks, LIBERO-100 and LIBERO-Plus demonstrate CoRe-VLA substantially improves task success across mainstream VLAs under camera shifts. For example, CoRe-VLA raises PI0.5's success rate from 13.3% to 83.3% at a 1.6m camera shift in real-robot environment.
Comments15 pages, 9 figures