强化视频推理分割以在分割前思考
Reinforcing Video Reasoning Segmentation to Think Before It Segments
AI总结:
Veason-R1通过组相对策略优化和链式思维初始化,提升视频推理分割的结构化推理能力,实现更准确的时空定位和鲁棒性。
AI中文摘要:
视频推理分割(VRS)旨在通过隐含的指令来界定视频中的所指对象,这些指令包含了人类意图和时间逻辑。先前的方法利用大型视觉语言模型(LVLMs)将物体语义编码为<SEG>标记以进行掩码预测。然而,这种范式在推理过程中存在解释性有限和性能不佳的问题,因为时空推理不足。受强化学习领域突破的启发,我们引入了Veason-R1,一个专门用于VRS的LVLM,强调分割中的结构化推理。Veason-R1通过组相对策略优化(GRPO)增强链式思维(CoT)初始化进行训练。首先,我们精心挑选高质量的CoT训练数据,以培养结构化的推理轨迹,连接视频层面的语义和帧层面的空间定位,产生监督微调模型Veason-SFT。随后,GRPO微调鼓励通过优化推理链高效探索推理空间。为此,我们引入了综合奖励机制,协同增强空间对齐和时间一致性,加强关键帧定位和细粒度定位。全面的实证评估表明,Veason-R1在多个基准上实现了最先进的性能,超越了先前的方法(例如,在ReVOS中+1.3 J &F,在ReasonVOS中+10.0 J &F),同时表现出对幻觉的鲁棒性(+8.8 R)。我们的代码和模型权重将在Veason-R1上提供。
英文摘要:
Video reasoning segmentation (VRS) endeavors to delineate referred objects in videos guided by implicit instructions that encapsulate human intent and temporal logic. Previous approaches leverage large vision language models (LVLMs) to encode object semantics into <SEG> tokens for mask prediction. However, this paradigm suffers from limited interpretability during inference and suboptimal performance due to inadequate spatiotemporal reasoning. Drawing inspiration from seminal breakthroughs in reinforcement learning, we introduce Veason-R1, a specialized LVLM for VRS that emphasizes structured reasoning in segmentation. Veason-R1 is trained through Group Relative Policy Optimization (GRPO) augmented with Chain-of-Thought (CoT) initialization. To begin with, we curate high-quality CoT training data to instill structured reasoning trajectories, bridging video-level semantics and frame-level spatial grounding, yielding the supervised fine-tuned model Veason-SFT. Subsequently, GRPO fine-tuning encourages efficient exploration of the reasoning space by optimizing reasoning chains. To this end, we incorporate a holistic reward mechanism that synergistically enhances spatial alignment and temporal consistency, bolstering keyframe localization and fine-grained grounding. Comprehensive empirical evaluations demonstrate that Veason-R1 achieves state-of-the-art performance on multiple benchmarks, surpassing prior art by significant margins (e.g., +1.3 J &F in ReVOS and +10.0 J &F in ReasonVOS), while exhibiting robustness to hallucinations (+8.8 R). Our code and model weights will be available at Veason-R1.