arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23138cs.ROcs.AIcs.CV

Pointing-VLA:面向视觉-语言-动作操控的类型化空间定位接口

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

Xiwen Chen, Zelin Li, Zhiruo Zhou, Huiming Chen, Chenwei Wang, Xiaojun Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

Pointing-VLA是基于Embodied-R1的类型化空间读出接口,可提升VLA模型的机器人操控性能,在多任务评估中实现SOTA表现,还能高效迁移至其他机器人系统并提升真实机器人自主成功率。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型通常通过自回归文本坐标或不透明的动作标记来暴露空间定位,这在多模态推理与机器人执行之间形成了脆弱的接口。我们提出了Pointing-VLA,一种构建于Embodied-R1之上的类型化隐藏状态空间读出机制。特定几何的头网络预测归一化点、对象功能定位(OFG)热图和视觉轨迹,而无需将几何序列化为文本。针对所评估的Bridge/WidowX和物理抓取-放置部署,显式执行契约将PICK(抓取)分配给源条件OFG,将PLACE(放置)分配给Pointing,提供与阶段对齐的直接空间目标。Pointing-VLA在Bridge/WidowX上实现了SOTA性能,在启用碰撞的CuRobo执行下,未进行Bridge特定微调的情况下,在评估的四项任务集上平均达到72.9%。Pointing和OFG在原生和跨数据集评估中表现出互补优势。OFG/接触读出可迁移至NORA-1.5,在保持或提升成功率的同时,将记录的控制器时间减少超过20倍;类型化头网络在共享外部套件上的速度也比Embodied-R1文本解码快6.68至6.90倍。当作为空间指导集成到π₀.₅动作策略中时,Pointing-VLA在三种视觉场景下将自主真实机器人成功率从52.7%提升至80.7%。这些结果确立了类型化空间读出是具身推理与机器人执行之间高效、可检查的接口。

英文摘要

Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. For the evaluated Bridge/WidowX and physical pick-place deployments, an explicit execution contract assigns PICK to source-conditioned OFG and PLACE to Pointing, providing direct stage-aligned spatial targets. Pointing-VLA achieves SOTA performance on Bridge/WidowX, averaging 72.9\% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Pointing and OFG show complementary strengths across native and cross-dataset evaluations. The OFG/contact readout transfers to NORA-1.5, preserving or improving success while reducing recorded controller time by more than 20$\times$; typed heads are also 6.68--6.90$\times$ faster than Embodied-R1 text decoding on a shared external suite. When integrated as spatial guidance for a $π_{0.5}$ action policy, Pointing-VLA raises autonomous real-robot success from 52.7\% to 80.7\% across three visual contexts. These results establish typed spatial readouts as an efficient, inspectable interface between embodied reasoning and robot execution.

↑