细粒度眼动视频时空定位
Fine-grained Spatiotemporal Grounding on Egocentric Videos
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出EgoMask基准和EgoMask-Train数据集,针对眼动视频细粒度时空定位的挑战,通过微调提升模型性能。
AI中文摘要:
时空视频定位旨在基于文本查询在视频中定位目标实体。尽管现有研究在非眼动视频方面取得了显著进展,但眼动设置仍相对研究较少,尽管其在增强现实和机器人等应用中日益重要。在本工作中,我们系统分析了眼动视频与非眼动视频之间的差异,揭示了关键挑战,如对象持续时间较短、轨迹较稀疏、对象尺寸较小和位置偏移较大。为解决这些挑战,我们引入了EgoMask,这是首个针对眼动视频细粒度时空定位的像素级基准。它通过我们提出的自动标注流程构建,该流程在短、中、长视频上标注指代表达和对象掩码。此外,我们创建了EgoMask-Train,一个大规模训练数据集以促进模型开发。实验表明,最先进的时空定位模型在我们的基准EgoMask上表现不佳,但通过在EgoMask-Train上微调可获得显著改进,同时在非眼动数据集上保持性能。我们的工作因此为推进眼动视频理解提供了必要的资源和见解。我们的代码可在https://github.com/LaVi-Lab/EgoMask获取。
英文摘要:
Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite its growing importance in applications such as augmented reality and robotics. In this work, we conduct a systematic analysis of the discrepancies between egocentric and exocentric videos, revealing key challenges such as shorter object durations, sparser trajectories, smaller object sizes, and larger positional shifts. To address these challenges, we introduce EgoMask, the first pixel-level benchmark for fine-grained spatiotemporal grounding in egocentric videos. It is constructed by our proposed automatic annotation pipeline, which annotates referring expressions and object masks across short-, medium-, and long-term videos. Additionally, we create EgoMask-Train, a large-scale training dataset to facilitate model development. Experiments demonstrate that the state-of-the-art spatiotemporal grounding models perform poorly on our benchmark EgoMask, but fine-tuning on EgoMask-Train yields significant improvements, while preserving performance on exocentric datasets. Our work thus provides essential resources and insights for advancing egocentric video understanding. Our code is available at https://github.com/LaVi-Lab/EgoMask .