事件驱动的刷新与循环记忆以减少指代视频对象分割中的过时定位
Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation
浏览论文内容
中文总结 AI 辅助
针对指代视频对象分割中因固定关键帧传播导致的过时定位问题,提出事件驱动刷新与循环记忆方法,在稳定变化点选择性重调用Sa2VA,以低刷新开销保持精度并减少误报。
中文摘要 AI 辅助
指代视频对象分割(RVOS)旨在为自然语言指定的对象生成像素级精确的掩码序列。Sa2VA将多模态大语言模型与SAM2相结合进行接地分割;然而,其推理通常从一组固定的初始关键帧中接地查询,然后依赖传播。在长视频或动态视频中,当对象组成发生变化(例如,干扰物进入或目标消失/重新出现)时,这可能导致过时定位和持续误报。我们提出事件驱动的刷新+循环记忆(EDRRM),一种增强方法,仅在稳定的变化点选择性重新调用Sa2VA。EDRRM使用从跟踪派生的线索(出生/死亡以及粗略组成/布局变化)计算的EMA平滑事件分数,并施加时间约束来触发刷新边界。循环记忆进一步通过CLIP相似性检索锚定帧,以在重新出现事件上重新调节模型。在Ref-DAVIS17、MeViS和ReVOS上的实验表明,相对于固定窗口和FrameDiff-SSIM基线,EDRRM实现了有竞争力的精度-效率权衡,在显著更低的平均刷新调用预算下保持相当或更优的J&F分数,并减少误报失败。端到端运行时分析进一步证实,相对于占主导地位的Sa2VA推理成本,跟踪、基于CLIP的循环匹配和可识别性门引入的开销仍然较小,从而验证了所提出流程的效率。
英文摘要
Referring Video Object Segmentation (RVOS) aims to produce a pixel-accurate mask sequence for an object specified by natural language. Sa2VA combines a multimodal large language model with SAM2 for grounded segmentation; however, its inference typically grounds the query from a small fixed set of initial keyframes and then relies on propagation. In long or dynamic videos, this can cause stale grounding and persistent false positives when the object composition changes (e.g., distractors enter or the target disappears/re-appears). We propose Event-Driven Refresh + Recurrence Memory (EDRRM), an enhancement that selectively re-invokes Sa2VA only at stable change points. EDRRM triggers refresh boundaries using an EMA-smoothed event score computed from tracking-derived cues (births/deaths and coarse composition/layout changes) with temporal constraints. A recurrence memory further retrieves anchor frames via CLIP similarity to re-condition the model on re-appearance events. Experiments on Ref-DAVIS17, MeViS, and ReVOS show that EDRRM achieves a competitive accuracy-efficiency trade-off relative to fixed-window and FrameDiff-SSIM baselines, maintaining comparable or superior J&F scores at substantially lower average refresh-call budgets and reducing false-positive failures. End-to-end runtime analysis further confirms that the overhead introduced by tracking, CLIP-based recurrence matching, and the identifiability gate remains modest relative to the dominant Sa2VA inference cost, thereby validating the efficiency of the proposed pipeline.
发表机构
- Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。