arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12220cs.CVcs.AI

SCOUT:通过结构化思维链与多目标过程奖励解锁增强型空间推理能力

SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward

Zile Zhou, Huining Yuan, Weichen Zhang, Xinlei Chen, Xiao-ping Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

SCOUT是结合结构化思维链与多目标过程奖励的视觉-语言模型,通过定制数据集训练,在空间推理任务上超越基线模型与GPT-4o,且具备跨图像、视频的泛化能力。

中文摘要 AI 辅助

现有视觉-语言模型(VLMs)在鲁棒空间推理方面存在关键瓶颈。近期强化学习(RL)方法试图通过可验证结果缩小这一差距,但它们在中间推理步骤间的信用分配效果较差。同时,结构化推理方法忽略了全面3D理解所需的关键深度感知。为解决这些挑战,我们提出SCOUT(利用过程监督强化学习训练的结构化思维链)。具体而言,我们设计了一个结构化思维链(CoT)框架,该框架明确建模3D环境感知以确保鲁棒的空间理解与推理。此外,我们引入一种新型强化学习算法,其具有多目标过程奖励和定制的优势估计方法,可在推理轨迹的不同片段间实现细粒度的信用分配。为支持我们的框架,我们开发了SCOUT-24k,这是一个通过定制流水线合成的结构化空间推理思维链数据集。大量评估表明,SCOUT-3B在通用空间基准和复杂空间推理任务上分别比基线模型提升了16.85%和6.3%。值得注意的是,我们更大规模的SCOUT-7B甚至以4.28%的优势超越了GPT-4o。此外,尽管仅在单图像上进行训练,SCOUT-7B仍展现出对多图像和视频场景的鲁棒域外泛化能力。这些实证结果表明,SCOUT是迈向下一代空间感知视觉-语言模型的关键一步。

英文摘要

Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning. Recent reinforcement learning (RL) methods aim to close this gap with verifiable outcomes, yet they suffer from poor credit assignment across intermediate reasoning steps. Concurrently, structured reasoning approaches overlook the critical depth perception necessary for comprehensive 3D understanding. To address these challenges, we propose SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training). Specifically, we design a structured Chain-of-Thought (CoT) framework that explicitly models 3D environmental perception to ensure robust spatial understanding and reasoning. Furthermore, we introduce a novel RL algorithm featuring multi-objective process rewards and a tailored advantage estimation method, facilitating fine-grained credit assignment across distinct segments of the reasoning trajectory. To support our framework, we develop SCOUT-24k, a structured spatial reasoning CoT dataset synthesized through a customized pipeline. Extensive evaluations demonstrate that SCOUT-3B improves upon baseline models by 16.85% and 6.3% on general spatial benchmarks and complex spatial reasoning tasks respectively. Notably, our larger SCOUT-7B even outperforms GPT-4o by a margin of 4.28%. Moreover, despite being trained exclusively on single image, SCOUT-7B exhibits robust out-of-domain generalization to multi-image and video scenarios. These empirical results render SCOUT as a critical step towards next generation of spatially-aware VLMs.

发表机构

  • Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑