空间潜在推理用于具身参考理解
Spatial Latent Reasoning for Embodied Reference Understanding
浏览论文内容
中文总结 AI 辅助
提出空间潜在推理框架,通过有序几何与视觉状态监督,提升指向手势视觉基础性能,在多个基准上显著优于现有方法。
中文摘要 AI 辅助
指向手势的视觉基础要求将手部几何形状与被指称对象的视觉身份和范围联系起来。连续潜在推理的一个核心挑战是如何将这些互补线索组织成有用的中间监督。我们提出了空间潜在推理(Spatial Latent Reasoning, SLR),一个围绕有序的几何和视觉状态序列来构建这种监督的框架。一个空间射线状态由指尖位置和指向方向监督,随后是四个与目标区域特征对齐的状态。为了构建视觉目标,我们引入了奇偶池化(parity pooling),它应用多相分组来在四个交错的空间支撑上平均区域令牌。所有状态在训练和推理期间循环生成;辅助标注仅在训练期间需要。在EgoPoint-Ground上,该框架在Qwen3.5-4B、Qwen2.5-VL-7B和Qwen3-VL-8B上分别比同骨干微调提高了2.8、17.5和21.1个百分点的mIoU,且在两个困难子集上均有改进。在YouRefIt上,它在IoU 0.5时达到77.6%的精确度,与所报告的最先进水平相比,在不同评估协议下具有5.2个百分点的数值优势。消融实验支持在标准和相似对象集上联合几何和视觉监督,并倾向于在标准集上使用奇偶池化而非三种替代池化算子。这些结果支持用于连续指向基础的任务结构化监督。我们将发布代码和支持材料。
英文摘要
Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object. A central challenge for continuous latent reasoning is how to organize these complementary cues into useful intermediate supervision. We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states. A spatial ray state is supervised by fingertip position and pointing direction, followed by four states aligned with target-region features. To construct the visual targets, we introduce parity pooling, which applies polyphase grouping to average region tokens on four interleaved spatial supports. All states are generated recurrently during training and inference; auxiliary annotations are required only during training. On EgoPoint-Ground, the framework improves mIoU over same-backbone supervised fine-tuning by 2.8, 17.5, and 21.1 percentage points on Qwen3.5-4B, Qwen2.5-VL-7B, and Qwen3-VL-8B, respectively, with improvements on both hard subsets. On YouRefIt, it achieves 77.6% precision at IoU 0.5, a numerical margin of 5.2 percentage points over the reported state of the art under differing evaluation protocols. Ablations support joint geometric and visual supervision on the standard and similar-object sets, and favor parity over three alternative pooling operators on the standard set. These results support task-structured supervision for continuous pointing grounding. We will release the code and supporting materials.
发表机构
- Tsinghua University(清华大学)
- Dalian University of Technology(大连理工大学)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。