AI 中文总结
针对多模态查询的指代视频分割,提出基于单一MLLM的OPERA智能体,通过时间与空间双轴渐进推理,在OmniAVS和Ref-AVS上达到最先进水平,并可零样本迁移。
AI 中文摘要
具有异构多模态查询(涵盖文本、音频和参考图像)的指代视频分割,既要求强大的跨模态理解能力,也要求精确的时空推理能力。我们提出了OPERA(全模态渐进式时空推理智能体),这是一个基于单一多模态大语言模型(MLLM)构建的统一推理智能体,通过三个专门阶段执行双轴渐进式推理。在时间轴上,时间推理智能体通过从粗到细的过滤来缩小帧搜索空间,以识别信息量最大的关键帧。在空间轴上,蒸馏智能体通过跨模态语义蒸馏确定要定位的目标,而由GRPO增强的定位智能体确定目标出现的位置,并通过密集掩码传播完成像素级输出。OPERA在OmniAVS和Ref-AVS上创下了新的最先进水平,并零样本迁移到标准的指代视频分割基准上。
英文摘要
Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent built on a single MLLM that performs dual-axis progressive reasoning via three specialized stages. Along the temporal axis, a Temporal Reasoning Agent narrows the frame search space through coarse-to-fine filtering to identify the most informative key frame. Along the spatial axis, a Distillation Agent establishes what to locate via cross-modal semantic distillation, and a Grounding Agent enhanced with GRPO determines where the target appears, with dense mask propagation completing the pixel-level output. OPERA sets a new state of the art on OmniAVS and Ref-AVS and transfers zero-shot to standard referring video segmentation benchmarks.
Comments17 pages