发表机构
Shanghai Jiao Tong University; Sun Yat-sen University; Fudan University(上海交通大学; 中山大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频分割中遮挡和重现导致的身份漂移问题,提出POSReasoner框架,通过持久状态和提议-验证推理决定状态更新,在冻结基础模型下提升长期VOS和VIS性能。
AI 中文摘要
视频分割模型通过跨帧携带实例信息来维持对象身份。然而,在长时间遮挡、重现或相似实例之间的交互情况下,不可靠的更新可能会覆盖有效历史并导致持续的身份漂移。我们提出了POSReasoner,一个可训练、即插即用的框架,它明确决定观测何时应改变对象的状态。每个持久状态记录身份、置信度和缺失历史。一个稀疏的状态-观测图支持“提议-验证”推理:使用对象历史、预测存在性和身份间的竞争来重新审视临时关联。验证后的决策决定是保留、更新、重新激活还是抑制每个状态,而一个学习到的门控控制写回记忆的证据。只有经过验证的转换才更新后续帧中使用的持久状态。POSReasoner使用标准视频标注并保持基础模型冻结,使其能够集成到多种VOS和VIS架构中。在长期VOS和VIS基准上的实验表明,与强基线相比有持续改进,在遮挡和对象重现情况下增益最大。
英文摘要
Video segmentation models maintain object identities by carrying instance information across frames. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an unreliable update can overwrite a valid history and cause persistent identity drift. We introduce POSReasoner, a trainable, plug-and-play framework that explicitly decides when an observation should change an object's state. Each persistent state records identity, confidence, and absence history. A sparse state-observation graph supports Propose-Verify reasoning: provisional associations are revisited using object history, predicted presence, and competition among identities. The verified decisions determine whether to retain, update, reactivate, or suppress each state, while a learned gate controls the evidence written back to memory. Only verified transitions update the persistent state used in subsequent frames. POSReasoner uses standard video annotations and keeps the base model frozen, enabling integration with diverse VOS and VIS architectures. Experiments across long-term VOS and VIS benchmarks show consistent improvements over strong baselines, with the largest gains under occlusion and object reappearance.
Comments19 pages, 6 figures