arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从推理失败到可组合的视频空间智能

From Reasoning Failures to Composable Video Spatial Intelligence

Pengzhan Sun, Junbin Xiao, Ramanathan Rajaraman, Shiu-hong Kao, Angela Yao

arXiv 2610.01999首次发表:更新:

发表机构

National University of Singapore; University of Science and Technology of China(新加坡国立大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过诊断视频空间推理中的四类错误,提出免训练的CROSS库,为视觉语言模型提供验证上下文或技能,在ReVSI和DSI-Bench上分别提升平均得分至60.2%和66.3%,证明显式空间约定可修复系统性推理失败。

AI 中文摘要

空间推理基准在多种任务上评估视觉语言模型,但任务级得分并不能揭示哪些底层能力导致了成功或失败。每项任务都需要恢复空间证据、表示几何结构并在此基础上进行推理。我们通过在一个共享模式与坐标约定下比较预测的和真实的空间上下文,将这些能力分离开来。这种比较揭示了四种反复出现的错误来源:感知不准确、空间上下文中的信息缺失、选择了错误的测量方式,以及参考系或位置与朝向跟踪中的错误。在这一诊断的指导下,我们开发了CROSS,一个免训练的库,包含类型化几何算子与空间技能,可在可用证据上运行以支持可靠的视频空间推理。由此产生的库为非编码的视觉语言模型提供经过验证的上下文,或为SpatialClaw智能体提供可调用的技能。我们在五个基准上评估了该方法。该方法将ReVSI上的平均得分从55.9%提升到60.2%,并将SpatialClaw在DSI-Bench上的结果从62.8%提升到66.3%。这些提升表明,显式处理空间约定可以在无需额外训练的情况下修复系统性的推理失败。

英文摘要

Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname{} on five benchmarks. \methodname{} raises the average score from 55.9\% to 60.2\% on ReVSI and improves the SpatialClaw result from 62.8\% to 66.3\% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑