arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpatialQuery:评估视觉语言模型中基于几何的多实例空间推理能力

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

Hai Nguyen, Tung Vu, Cong Tran

arXiv 2608.01709首次发表:更新:

发表机构

Posts and Telecommunications Institute of Technology(邮电技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉语言模型的多实例空间推理缺陷,推出SPATIALQUERY框架及含百万对问答的SPATIALQUERY-1M基准,结合UA-CoT提示,使Qwen3-VL-8B在空间推理任务中表现优于多种模型。

AI 中文摘要

视觉语言模型(VLMs)在语义理解上表现出色,但在度量空间推理方面仍不可靠,尤其是当查询需要比较同一类别的多个实例时。我们通过最近实例距离查询(CIDQ)研究该问题,要求模型识别距离唯一参考对象最近的可见候选对象,并估算它们与重力对齐的地面平面距离。我们推出SPATIALQUERY,这是一种无需训练的框架,用于从单张RGB图像进行CIDQ推理,同时推出SPATIALQUERY-1M,这是一个包含超过100万对仅RGB图像的问答对的基准,来自200个室内场景。SPATIALQUERY通过场景立方体化恢复实例级度量几何,并将其转换为标准鸟瞰图,将对象表示为大小一致、类别编码的块,以强调它们在地面平面上的相对位置。我们进一步提出不确定性感知思维链(UA-CoT)提示,将几何推导的每个实例的不确定性纳入VLM推理过程。无需特定任务微调或架构修改,结合Qwen3-VL-8B的SPATIALQUERY实现了0.259米的Floor-MAE、90.5%的Unc-Acc@0.3米以及84.18%的邻近决策准确率,优于微调的空间专家、通用VLMs和闭源前沿模型。代码、基准资源和交互式演示可在指定URL获取。

英文摘要

Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require comparing multiple instances of the same object category. We study this problem through the Closest-Instance Distance Query (CIDQ), where a model must identify the nearest visible candidate to a unique reference object and estimate their gravity-aligned floor-plane distance. We introduce SPATIALQUERY, a training- free framework for CIDQ reasoning from a single RGB image, together with SPATIALQUERY-1M, a benchmark containing over one million RGB-only question-answer pairs from 200 indoor scenes. SPATIALQUERY recovers instance-level metric geometry and transforms it into a canonical Bird's-Eye View through Scene Cubifying, which represents objects as uniformly sized, category-coded blocks to emphasize their relative floor- plane locations. We further propose Uncertainty-Aware Chain-of-Thought (UA-CoT) prompting, which incorporates geometry- derived per-instance uncertainty into the VLM reasoning process. Without task-specific fine-tuning or architectural modification, SPATIALQUERY with Qwen3-VL-8B achieves a Floor-MAE of 0.259 m, an Unc-Acc@0.3 m of 90.5%, and a proximity-decision accuracy of 84.18%, outperforming fine-tuned spatial specialists, general-purpose VLMs, and closed-source frontier models. Code, benchmark resources, and an interactive demo are available at https://namhai1810.github.io/SpatialQuery/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑