arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GroundAnything:以闪电速度调和并行解码与精确视觉定位

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou, Ping Luo, Shiyu Huang

arXiv 2609.39600首次发表:更新:

发表机构

XPeng Inc.; Peking University; The University of Hong Kong; University of California, Berkeley; Princeton University; National University of Singapore; HKUST (GZ)(小鹏汽车; 北京大学; 香港大学; 加州大学伯克利分校; 普林斯顿大学; 新加坡国立大学; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出GroundAnything,一个4B参数的视觉定位基础模型,通过分块去噪的并行解码实现快速且精确的定位,在30个基准上超越先前最先进水平,并支持自我推测加速。

AI 中文摘要

自回归(AR)定位模型将空间预测序列化,引入了顺序延迟,并对输出令牌施加了因果顺序。我们将定位视为视觉证据提取:对象、位置和空间关系都受图像和查询的联合约束,但它们的依赖关系并不意味着固有的从左到右的生成顺序。这一区别使得双向扩散成为自然的选择,允许空间假设并行出现,并通过迭代去噪进行联合细化。我们引入了GroundAnything,一个4B参数的定位基础模型,通过分块去噪调和了快速并行解码与精确定位。训练结合了来自公共数据集和专用数据引擎的定位预训练、直接AR到扩散的转换并联合AR和扩散目标、监督微调以及基于GRPO的强化后训练。在30个定位基准上,我们的自回归变体GroundAnything-VLM在相似规模模型中建立了新的整体最先进水平,达到72.42%,与GPT-6 Astra(71.35%)保持竞争力。通过熵引导解码,GroundAnything也超越了该规模下的先前最先进水平,平均61.75%,而基于快速MTP的LocateAnything模型为53.32%。我们进一步探索了解码策略,表明可选的自我推测模式相比AR对应模型实现了4.51倍的加速,但COCO F1mIoU下降了0.74个百分点。基础设施实验表明,渐进式推理优化将并行解码转化为实际加速。这些支持在延迟敏感的现实世界系统中进行高效的视觉定位。

英文摘要

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a $4.51\times$ speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.

Comments61 pages, including supplementary material. Project page: https://groundingpi.github.io/groundanything/ Code: [https://github.com/groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything) Model: https://huggingface.co/GroundingPI/GroundAnything, https://huggingface.co/GroundingPI/GroundAnything-VLM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑