发表机构
School of Artificial Intelligence and Computer Science, Jiangnan University; Centre for Vision, Speech and Signal Processing, University of Surrey(江南大学人工智能与计算机科学学院; 萨里大学视觉、语音和信号处理中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对引用视频对象分割问题,提出无训练、反馈驱动的ReflexTrack智能体。通过掩码引导空间细化和视频级掩码反射实现空间与时间层面反馈,提升预测可靠性,在Ref-VPS和ReasonVOS上取得较好分数。
AI 中文摘要
引用视频对象分割(RVOS)要求在整个视频中分割由自然语言指定的目标。最近的智能体方法将多模态大语言模型与可提示的分割模型相结合,以在无需特定任务训练的情况下执行RVOS。然而,大多数流程依赖于一次性空间定位,然后进行掩码传播,使得初始提示和时间预测在很大程度上未经验证。我们引入了ReflexTrack,一种无训练、反馈驱动的智能体,它在空间和时间层面上都闭合了这个循环。掩码引导的空间细化评估当前关键帧提示诱导的掩码,并迭代更新边界框以及正、负点,产生更可靠的初始化。视频级掩码反射评估完整的掩码序列,定位不可靠区间,选择互补的修复关键帧,并通过掩码引导的重新传播生成候选预测。只有提供经过验证的改进的候选预测才用于更新受影响的区间,同时保留其他可靠的预测。所有组件在推理期间保持冻结状态。ReflexTrack在Ref-VPS上的总体Q分数为69.7,在ReasonVOS上的J&F分数为67.2。这些结果表明,预测级反馈显著提高了无训练RVOS的可靠性。
英文摘要
Referring video object segmentation (RVOS) requires segmenting a target specified by natural language throughout a video. Recent agentic approaches combine multimodal large language models with promptable segmentation models to perform RVOS without task-specific training. However, most pipelines rely on one-shot spatial grounding followed by mask propagation, leaving both the initial prompts and temporal predictions largely unverified. We introduce ReflexTrack, a training-free, feedback-driven agent that closes this loop at both spatial and temporal levels. Mask-guided Spatial Refinement evaluates the mask induced by the current keyframe prompt and iteratively updates the bounding box together with positive and negative points, yielding a more reliable initialization. Video-level Mask Reflection assesses the complete mask sequence, localizes unreliable intervals, selects complementary repair keyframes, and generates candidate predictions through mask-guided re-propagation. Only candidates that provide a verified improvement are used to update the affected intervals, preserving reliable predictions elsewhere. All components remain frozen during inference. ReflexTrack achieves an overall $\mathcal{Q}$ score of $69.7$ on Ref-VPS and a $\mathcal{J}\&\mathcal{F}$ score of $67.2$ on ReasonVOS. These results demonstrate that prediction-level feedback substantially improves the reliability of training-free RVOS.