arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14466cs.CVcs.RO

DynEoMT:从在线分割查询中学习物体动态性

DynEoMT: Learning Object Dynamicity from Online Segmentation Queries

  • Université Paris-Saclay(巴黎-萨克雷大学)
  • CEA(法国原子能委员会)
  • List(List研究所)
  • ENSTA Paris(巴黎高科先进技术学校)
  • Institut Polytechnique de Paris(巴黎综合理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Calvin Galagain, Martyna Poreba, François Goulette

AI总结:

DynEoMT是一个在线框架,通过传播查询为视频分割区域预测动态性,无需光流或深度,在多个基准上取得高平衡准确率,同时保持分割性能。

AI中文摘要:

视频分割模型能够随时间识别和跟踪物体,但它们并不指示每个分割区域是否独立于观察相机移动。这种动态性属性无法仅从语义中推断,并且会受到相机自身运动的干扰。我们引入了DynEoMT,一个在线框架,为基于查询的视频分割增加了区域级动态性预测。它同时产生原始的分割输出以及每个预测区域的动态或静态状态。在推理时,DynEoMT仅使用当前帧和传播的查询,不使用光流、深度、相机位姿、先前的RGB帧或特征图。由于现有的视频分割基准未标注此属性,我们还引入了一个类无关的离线监督流程,使用相机补偿的光流和置信度感知的时间滤波。在VIPSeg、OVIS、YouTube-VIS 2022和VSPW上,DynEoMT分别实现了84.3、68.0、68.6和87.6的平衡准确率,同时基本保持了分割性能。这些结果表明,分割区域的动态性可以从传播的查询中学习,从而在推理时无需专门的运动处理流程即可进行在线预测。完整代码将作为开源发布,以实现方法和实验的完全复现。

英文摘要:

Video segmentation models recognize and track objects over time, but they do not indicate whether each segmented region moves independently of the observing camera. This dynamicity attribute cannot be inferred from semantics alone and is confounded by camera ego-motion. We introduce \method, an online framework that augments query-based video segmentation with region-level dynamicity prediction. It jointly produces the original segmentation outputs and a dynamic or static state for each predicted region. At inference, DynEoMT uses only the current frame and propagated queries, without optical flow, depth, camera pose, previous RGB frames, or feature maps. Because established video segmentation benchmarks do not annotate this attribute, we also introduce a class-agnostic offline supervision pipeline using camera-compensated optical flow and confidence-aware temporal filtering. Across VIPSeg, OVIS, YouTube-VIS 2022, and VSPW, DynEoMT achieves balanced accuracies of 84.3, 68.0, 68.6, and 87.6, respectively, while largely preserving segmentation performance. These results show that segmentation-region dynamicity can be learned from propagated queries, enabling its online prediction without a dedicated motion-processing pipeline at inference. The complete code will be released as open source to enable full reproduction of the method and experiments.

补充信息

↑