发表机构
XPeng Inc.; Peking University; The University of Hong Kong; University of California, Berkeley; Princeton University; National University of Singapore; Tsinghua University; HKUST (GZ)(小鹏汽车; 北京大学; 香港大学; 加州大学伯克利分校; 普林斯顿大学; 新加坡国立大学; 清华大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GroundingPI,一个4B参数的接地基础模型,通过共享词汇量化生成点和框,在34个基准上超越44个基线,并显著提升机器人操作与自动驾驶性能,验证了专用感知预训练对物理智能的价值。
AI 中文摘要
精确的接地至关重要。它指定了哪个对象是目标以及该对象的位置,即使在杂乱环境中且对象很小的情况下也是如此,并且它必须足够快以用于闭环控制。然而,视觉-语言-动作(VLA)模型和世界-动作模型(WAMs)从通用视觉-语言和视频生成骨干网络中获取感知,这些骨干网络在这些场景中仍然失败。我们引入了GroundingPI,一个4B参数的接地基础模型,将点和框生成为共享词汇表中的量化坐标。训练结合了多模态和空间预训练、监督微调以及使用GRPO的强化学习,利用来自公共数据集和专用数据引擎的监督。在跨越11种感知能力的34个接地基准上,与44个基线相比,GroundingPI建立了新的最先进水平,平均达到73.68%,高于更大的GPT-6 Astra(71.54%)。作为下游视觉骨干,GroundingPI提升了机器人操作和自动驾驶的性能。在RoboTwin 2.0上,它在所有四种分布外设置中均优于我们评估的每个主流骨干,相对于最强骨干最高提升24.8%。在RoboCasa-GR1上,使用50%演示训练的GroundingPI优于使用75%演示训练的基线。在nuScenes上,作为视觉骨干,GroundingPI实现了平均开环L2误差0.296米。我们系统地分析了GroundingPI在规模和数据组成方面的预训练。随着预训练规模的扩大,下游自动驾驶和机器人操作性能得到提升。对这11种感知能力的数据配方分析表明,密集接地对两者都有显著益处,而OCR作为感知学习的催化剂具有潜力。这些结果支持将接地作为感知基础,以及专门的感知预训练作为物理智能基础模型的一个有前景的方向。
英文摘要
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.
Comments64 pages, including supplementary material. Project page: https://groundingpi.github.io/ Code: https://github.com/groundingpi/GroundingPI Model: https://huggingface.co/GroundingPI/GroundingPI