arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04185cs.CV

移动相机视频中的指代多目标跟踪:基于全局运动补偿

Referring Multi-Object Tracking in Moving-Camera Videos via Global Motion Compensation

Hsin-Chen Pai, Jyun-Kai Wang, Yi-Cheng Peng, Wei-Ta Chu

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对移动相机视频中的指代多目标跟踪,提出通过估计相机运动提取残差运动并与查询表达式匹配,经后期融合提升RMOT性能。

中文摘要 AI 辅助

指代多目标跟踪(RMOT)以视频和语言表达式为输入,并跟踪所有被指代的对象。许多跟踪需求涉及对象的运动方式而非外观。然而,在由移动相机拍摄的视频中,停放的车辆可能看起来在移动,而行驶中的车辆可能显示很小的位移。现有的RMOT方法将运动与文本相关联,但并未显式去除相机引起的运动。在本文中,我们提出通过估计驾驶员视角视频中的相机运动来提取跨帧的残差运动,并将运动特征与查询表达式进行比较。我们考虑运动匹配程度,并通过后期融合将其与RMOT方法的预测结果相结合。在评估中,我们验证了将运动补偿模块作为插件在不同RMOT主机上的性能增益。

英文摘要

Referring multi-object tracking (RMOT) takes a video and a language expression as input and tracks all referred objects. Many tracking requirements involve how an object moves rather than how it appears. However, in a video captured by a moving camera, a parked vehicle may appear to move, while a moving vehicle may show little displacement. Existing RMOT methods relate motion with text but do not explicitly remove camera-induced motion. In this paper, we propose extracting residual motion across frames by estimating camera motion in driver-view videos and compare motion characteristics with the query expression. We consider the motion-matching extent and integrate it with the RMOT method's prediction result through late fusion. In the evaluation, we verify the performance gain of taking the motion compensation module as a plug-in across different RMOT hosts.

发表机构

  • National Cheng Kung University(国立成功大学)

机构由 AI 辅助整理,请以论文原文为准。

↑