发表机构
JD Explore Academy; School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen; Shenzhen Institute of Artificial Intelligence and Robotics for Society; South China University of Technology; Sun Yat-sen University; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(京东探索研究院; 香港中文大学(深圳)理工学院; 深圳市人工智能与机器人研究院; 华南理工大学; 中山大学; 香港中文大学(深圳)人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出CometVLA,通过构建与机器人动作对齐的具身物理VQA数据及引入GAP token,在具身数据金字塔上协同训练,提升VLA模型的物理理解能力,显著优于基线模型并验证了物理预训练对下游操作的益处。
AI 中文摘要
视觉-语言-动作(VLA)模型在需要物理常识的操作任务中仍存在脆弱性。当前的物理视觉问答(VQA)数据通常是非具身的,与机器人动作领域不匹配;自我中心视频仅用作辅助预训练。目前尚不清楚视觉语言模型(VLM)物理理解的提升是否真的有益于下游动作生成。因此,我们提出CometVLA以缩小这一差距。我们构建了CometData和CometBench,这是与机器人动作数据及具身性严格对齐的具身物理VQA语料库和基准。我们引入了全局动作先验(GAP)token,这是一种紧凑的可学习瓶颈,可分离与任务无关的运动规律,使动作头能利用物理常识而不破坏预训练的VLM主干。我们在跨越遥操作、仿真、自我中心轨迹和VQA层的具身数据金字塔上协同训练CometVLA。在真实世界操作任务和RoboTwin仿真中,CometVLA始终优于强大的VLA基线。相关性分析显示,VLM在CometBench上的性能越强,VLA的成功率越高。结果表明,物理理解预训练确实有益于下游操作。
英文摘要
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understanding actually benefits downstream action generation. Therefore, we present CometVLA to close this gap. We construct CometData and CometBench, an embodied physical VQA corpus and benchmark strictly aligned with the robot's action data and embodiment. We introduce Global Action Prior (GAP) tokens, a compact learnable bottleneck that isolates task-agnostic motion regularities and lets the action head consume physical commonsense without corrupting the pre-trained VLM backbone. We co-train CometVLA across the embodied data pyramid, spanning teleoperation, simulation, egocentric trajectories, and VQA layers. On real-world manipulation tasks and RoboTwin simulation, CometVLA consistently outperforms strong VLA baselines. Correlation analysis shows that stronger VLM performance on CometBench indicates higher VLA success rates. Results demonstrate that physical understanding pre-training genuinely benefits downstream manipulation.