arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OPERA:用于指代视频分割的统一全模态渐进式时空推理智能体

OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation

Jingchen Ni, Yuji Wang, Shannan Yan, Haoru Li, Sitong Chen, Chun Yuan

arXiv 2609.33338首次发表:更新:

AI 中文总结

针对多模态查询的指代视频分割,提出基于单一MLLM的OPERA智能体,通过时间与空间双轴渐进推理,在OmniAVS和Ref-AVS上达到最先进水平,并可零样本迁移。

AI 中文摘要

具有异构多模态查询(涵盖文本、音频和参考图像)的指代视频分割,既要求强大的跨模态理解能力,也要求精确的时空推理能力。我们提出了OPERA(全模态渐进式时空推理智能体),这是一个基于单一多模态大语言模型(MLLM)构建的统一推理智能体,通过三个专门阶段执行双轴渐进式推理。在时间轴上,时间推理智能体通过从粗到细的过滤来缩小帧搜索空间,以识别信息量最大的关键帧。在空间轴上,蒸馏智能体通过跨模态语义蒸馏确定要定位的目标,而由GRPO增强的定位智能体确定目标出现的位置,并通过密集掩码传播完成像素级输出。OPERA在OmniAVS和Ref-AVS上创下了新的最先进水平,并零样本迁移到标准的指代视频分割基准上。

英文摘要

Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent built on a single MLLM that performs dual-axis progressive reasoning via three specialized stages. Along the temporal axis, a Temporal Reasoning Agent narrows the frame search space through coarse-to-fine filtering to identify the most informative key frame. Along the spatial axis, a Distillation Agent establishes what to locate via cross-modal semantic distillation, and a Grounding Agent enhanced with GRPO determines where the target appears, with dense mask propagation completing the pixel-level output. OPERA sets a new state of the art on OmniAVS and Ref-AVS and transfers zero-shot to standard referring video segmentation benchmarks.

Comments17 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑