发表机构
XPeng Inc.; Peking University; The University of Hong Kong; University of California, Berkeley; Princeton University; National University of Singapore; HKUST (GZ)(小鹏汽车; 北京大学; 香港大学; 加州大学伯克利分校; 普林斯顿大学; 新加坡国立大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GroundAnything,一个4B参数的视觉定位基础模型,通过分块去噪的并行解码实现快速且精确的定位,在30个基准上超越先前最先进水平,并支持自我推测加速。
AI 中文摘要
自回归(AR)定位模型将空间预测序列化,引入了顺序延迟,并对输出令牌施加了因果顺序。我们将定位视为视觉证据提取:对象、位置和空间关系都受图像和查询的联合约束,但它们的依赖关系并不意味着固有的从左到右的生成顺序。这一区别使得双向扩散成为自然的选择,允许空间假设并行出现,并通过迭代去噪进行联合细化。我们引入了GroundAnything,一个4B参数的定位基础模型,通过分块去噪调和了快速并行解码与精确定位。训练结合了来自公共数据集和专用数据引擎的定位预训练、直接AR到扩散的转换并联合AR和扩散目标、监督微调以及基于GRPO的强化后训练。在30个定位基准上,我们的自回归变体GroundAnything-VLM在相似规模模型中建立了新的整体最先进水平,达到72.42%,与GPT-6 Astra(71.35%)保持竞争力。通过熵引导解码,GroundAnything也超越了该规模下的先前最先进水平,平均61.75%,而基于快速MTP的LocateAnything模型为53.32%。我们进一步探索了解码策略,表明可选的自我推测模式相比AR对应模型实现了4.51倍的加速,但COCO F1mIoU下降了0.74个百分点。基础设施实验表明,渐进式推理优化将并行解码转化为实际加速。这些支持在延迟敏感的现实世界系统中进行高效的视觉定位。
英文摘要
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a $4.51\times$ speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.
Comments61 pages, including supplementary material. Project page: https://groundingpi.github.io/groundanything/ Code: [https://github.com/groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything) Model: https://huggingface.co/GroundingPI/GroundAnything, https://huggingface.co/GroundingPI/GroundAnything-VLM