SpatialCLI:先借助空间工具进行推理,再脱离工具进行推理
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
浏览论文内容
中文总结 AI 辅助
本研究提出SpatialCLI框架,通过三个阶段让视觉语言模型先借助专用空间工具推理,再内化感知能力,在MindCube上大幅提升了Qwen3-VL-8B-Instruct的性能,表现优于对比模型。
中文摘要 AI 辅助
视觉语言模型(VLMs)越来越多地被用于具身智能体,以解释视觉输入、推理空间关系并基于此做出任务级决策。然而,仍存在根本性的能力不匹配:通用VLMs可对整体任务进行推理,但往往会遗漏决定成功与否的视觉细节;而专用视觉模型能捕捉这些细节,却无法将其转化为任务级决策。在本研究中,我们提出SpatialCLI框架,该框架教导VLMs借助空间工具进行推理,并逐步内化这些工具提供的专用感知能力。SpatialCLI分为三个阶段:(1)调用阶段将专用视觉模型作为空间工具,以增强VLM的感知能力;(2)学习阶段采用冷启动监督微调(Cold-Start SFT)和智能体强化学习(agentic RL)来提升工具使用能力;(3)内化阶段将成功的工具使用轨迹进行语言表述,以内化专用感知能力。我们还推出了SpatialCLI-Bench,这是一个包含516个示例的基准,用于定位、分割、深度和姿态方面的组合感知。在MindCube上,SpatialCLI借助工具将Qwen3-VL-8B-Instruct的性能从29.3%提升至84.6%,超过了借助工具的GPT-5.6 Sol(72.1%),在工具内化后,脱离工具仍能保持73.8%的性能。
英文摘要
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM's perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.