Spatial-OPSD:通过无标签自蒸馏实现自我改进的空间推理
Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation
查看机构详情
- Tsinghua University(清华大学)
- LMMs-Lab
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出Spatial-OPSD,一种无标签自蒸馏框架,利用深度和3D重建等空间先验作为教师信号,通过递归训练提升VLM空间推理能力,在多个基准上达到开源最优。
中文摘要 AI 辅助
视觉语言模型(VLMs)越来越多地应用于具身和空间接地环境中,在这些环境中,准确理解深度、视角和三维关系至关重要。然而,改进空间推理通常依赖于真实答案、基于答案的奖励或其他形式的任务特定监督。我们引入了Spatial-OPSD,一种无标签的自我改进框架,它转而利用从感知和重建工具中自然获得的空间结构。在训练期间,特权教师接收自动可获得的空间先验,如深度、重建的三维关系和相机几何,而学生仅观察原始的视觉语言输入。在学生自身采样的轨迹上,教师提供密集的令牌级监督,使学生能够在没有真实答案标签或推理时特权信息的情况下内化空间知识。为了将这种监督扩展到单轮之外,我们采用了一种逐轮递归训练方案:每轮内教师保持冻结以提供稳定的学习目标,改进后的学生在下一轮中初始化教师和学生,其中特权空间先验重新建立信息丰富的师生不对称性。这实现了重复的自我改进,同时避免了优化过程中教师快速移动的问题。在四个VLM家族中,单轮Spatial-OPSD持续提高了五个基准的平均值,而三轮进一步将强大的空间专门模型推向了开源前沿,在开放模型中取得了最高平均值,并在五个空间推理基准中的三个上取得了最佳结果。我们的代码可在以下网址获取:此https URL。
英文摘要
Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher--student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at https://github.com/vermouth599/Spatial-OPSD.