arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24598cs.CV

QueenVIS:通过查询增强对视频实例分割仅图像训练的重新思考

QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

Arian Kheirandish, Fardin Ayar, Ehsan Javanmardi, Manabu Tsukada, Mahdi Javanmardi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对视频实例分割中仅图像训练方法的不足,提出QueenVIS框架,通过单帧训练时用辅助头丰富查询,推理时丢弃辅助头,以无训练方案保持时间身份,实验证明该方法有效提升分割精度,凸显增强查询稳定性对VIS的重要性。

中文摘要 AI 辅助

视频实例分割(VIS)要求模型在帧间检测、分割并跟踪对象身份,多数方法通过视频级监督强化时间一致性。以MinVIS为代表的仅图像训练方法对此提出挑战,将帧视为独立图像,仅在推理时关联实例,无需视频训练即可达到有竞争力的VIS。然而该领域已转向更复杂的视频训练跟踪器,依赖昂贵的身份一致注释,使仅图像方向未得到充分探索。诊断分析表明对象查询质量是瓶颈:仅在帧内定位对象训练的查询在帧间会漂移,破坏关联。QueenVIS引入以查询为中心的框架来强化仅图像训练的VIS。单帧训练期间,用两个辅助头丰富Mask2Former查询:一个特征预测损失使每个查询与其实例的池化主干描述符对齐,一个中心预测损失注入空间结构。推理时丢弃这两个头,不增加参数,通过无训练的查询传播和内存库方案保持时间身份。在YouTube-VIS和OVIS上,使用ResNet-50主干的QueenVIS比MinVIS有改进,在YouTube-VIS上提升高达+6.7 AP,在OVIS上提升+4.8 AP,在长序列YouTube-VIS分割上提升+10.3 AP。QueenVIS在YouTube-VIS上达到50.9 AP,与近期视频监督的最先进方法竞争,训练期间不处理单个视频片段。研究结果表明,增强对象查询的判别力和时间稳定性是VIS一个重要但未充分探索的方向。

英文摘要

Video instance segmentation (VIS) requires models to detect, segment, and track object identities across frames, and most methods enforce temporal consistency through video-level supervision. Image-only training approaches, with MinVIS as one prominent example, have challenged this assumption, reaching competitive VIS without video training by treating frames as independent images and associating instances only at inference. The field has nonetheless moved toward ever more elaborate video-trained trackers, which depend on costly identity-consistent annotations, leaving the image-only direction under-explored. A diagnostic analysis identifies object query quality as the bottleneck: queries trained only to localize objects within a frame drift apart across frames and destabilize association. QueenVIS introduces a query-centric framework for strengthening image-trained VIS. During single-frame training, we enrich Mask2Former queries with two auxiliary heads: a feature-prediction loss that aligns each query with the pooled backbone descriptor of its instance, and a center-prediction loss that injects spatial structure. Both heads are discarded at inference, adding zero parameters, and temporal identity is maintained by a training-free query-propagation and memory-bank scheme. On YouTube-VIS and OVIS with a ResNet-50 backbone, QueenVIS improves over MinVIS, up to +6.7 AP on YouTube-VIS, +4.8 AP on OVIS, and +10.3 AP on the long-sequence YouTube-VIS split. QueenVIS achieves 50.9 AP on YouTube-VIS and remains competitive with recent video-supervised state-of-the-art, without processing a single video clip during training. Our findings suggest that strengthening the discriminative power and temporal stability of object queries is an important, underexplored axis for VIS. Code and models: https://github.com/ArianKheir/QueenVIS

发表机构

  • Amirkabir University of Technology (AUT)(伊朗阿米尔卡比尔理工大学)
  • The University of Tokyo(日本东京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑