解锁以外部视角视频-语言数据用于第一人称视频表示学习
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
浏览论文内容
中文总结 AI 辅助
本文提出 EMBED 方法,通过视觉和语言风格迁移将外部视角视频-语言数据转换为第一人称训练数据,并在多个第一人称下游任务上取得最先进结果,同时具备跨视角泛化能力。
中文摘要 AI 辅助
我们提出 EMBED(Egocentric Models Built with Exocentric Data,基于外部视角数据构建的第一人称模型),这是一种旨在将外部视角视频-语言数据转化为第一人称视频表示学习的方法。大规模外部视角数据覆盖多样化的活动,对第一人称学习具有巨大潜力,但第一人称数据与外部视角数据之间的固有差异使得无缝地利用一种视角为另一种视角服务面临挑战。第一人称视频主要呈现近距离的手-物交互,而外部视角视频则提供更广阔的人类活动视角。此外,第一人称数据集中的叙述通常更以动作为中心,并与视觉内容紧密关联,这与外部视角数据集中常见的叙述风格不同。为解决这些挑战,我们采用一种数据转换框架来适配外部视角数据以用于第一人称训练,重点在于识别强调手-物交互的特定视频片段,并转换叙述风格以符合第一人称视角。通过同时应用视觉和语言风格迁移,我们的框架从外部视角视频-语言数据中构建了一个新的第一人称数据集。通过广泛评估,我们证明了 EMBED 的有效性,在多个第一人称下游任务上取得了最先进的结果,包括在零样本设置下 Epic-Kitchens-100 多实例检索上绝对提升 4.7%,在 EGTEA 分类基准上绝对提升 6.2%。此外,EMBED 使第一人称视频-语言模型能够在外部视角任务中表现出具有竞争力的性能。最后,我们展示了 EMBED 在多个外部视角数据集上的应用,表现出在不同外部视角数据集上应用时的强大泛化能力。
英文摘要
We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse activities with significant potential for egocentric learning, but inherent disparities between egocentric and exocentric data pose challenges in utilizing one view for the other seamlessly. Egocentric videos predominantly feature close-up hand-object interactions, whereas exocentric videos offer a broader perspective on human activities. Additionally, narratives in egocentric datasets are typically more action-centric and closely linked with the visual content, in contrast to the narrative styles found in exocentric datasets. To address these challenges, we employ a data transformation framework to adapt exocentric data for egocentric training, focusing on identifying specific video clips that emphasize hand-object interactions and transforming narration styles to align with egocentric perspectives. By applying both vision and language style transfer, our framework creates a new egocentric dataset derived from exocentric video-language data. Through extensive evaluations, we demonstrate the effectiveness of EMBED, achieving state-of-the-art results across various egocentric downstream tasks, including an absolute improvement of 4.7% on the Epic-Kitchens-100 multi-instance retrieval and 6.2% on the EGTEA classification benchmarks in zero-shot settings. Furthermore, EMBED enables egocentric video-language models to perform competitively in exocentric tasks. Finally, we showcase EMBED's application across various exocentric datasets, exhibiting strong generalization capabilities when applied to different exocentric datasets.
发表机构
- FAIR at Meta(Meta FAIR研究院)
- UCLA(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。