发表机构
Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对长期视频目标分割的误差累积问题,提出无训练即插即用的SAM2Dual,通过双内存设计与文本感知内存提升长视频鲁棒性,在MOSEv2、LVOSv2基准上取得性能提升。
AI 中文摘要
长期视频目标分割(VOS)因长时间遮挡、目标重现及场景变化下的误差累积仍具挑战性。尽管SAM2具备强大的零样本性能,但其流式内存会因近期不可靠预测主导内存状态,在长时序中放大漂移。本文提出SAM2Dual,这是一种无训练、即插即用的推理时增强方案,无需更新模型权重即可提升长视频鲁棒性。SAM2Dual引入双内存设计,明确区分用于快速局部适应的短期内存,以及通过间隔采样构建以保留全局身份线索的长期内存,并通过门控融合策略将二者结合。此外,本文提出文本感知内存(TAM),其从早期帧提取紧凑的词级线索,并利用文本嵌入基于语义兼容性重新加权内存贡献,在视觉证据薄弱或模糊时支持身份保留。在长期基准测试中,SAM2Dual持续提升长视频稳定性,使MOSEv2上的J&F从49.33提升至50.65,并在LVOSv2上实现持续增益。
英文摘要
Long-term video object segmentation (VOS) remains challenging due to error accumulation under extended occlusions, re-appearance, and scene changes. Although SAM2 provides strong zero-shot performance, its streaming memory can amplify drift over long horizons when recent, unreliable predictions dominate the memory state. We propose SAM2Dual, a training-free, plug-and-play inference-time enhancement that improves long-video robustness without updating model weights. SAM2Dual introduces a Dual Memory design that explicitly separates (i) short-term memory for rapid local adaptation and (ii) long-term memory built via interval-based sampling to preserve global identity cues, combined through a gated fusion strategy. In addition, we present Text-Aware Memory (TAM), which extracts a compact word-level cue from early frames and uses text embeddings to reweight memory contributions based on semantic compatibility, supporting identity preservation when visual evidence becomes weak or ambiguous. Across long-term benchmarks, SAM2Dual consistently improves stability on long videos, raising J&F from 49.33 to 50.65 on MOSEv2 and achieving consistent gains on LVOSv2.