Exo2Ego:基于外部知识引导的多模态大语言模型用于第一人称视频理解
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
浏览论文内容
中文总结 AI 辅助
Exo2Ego通过迁移学习提升内向视频理解能力,利用外向知识增强模型性能。
中文摘要 AI 辅助
人工智能个人助手通过机器人或可穿戴设备部署,需要具身理解才能有效与人类协作。然而,当前的多模态大语言模型(MLLMs)主要关注第三人称(外向)视觉,忽视了第一人称(内向)视频的独特挑战。此外,高采集成本限制了数据量,影响了MLLM性能。为解决这些挑战,我们提出了学习外向与内向领域之间的映射,利用现有MLLM中的广泛外向知识来增强内向视频理解。为此,我们引入了Ego-ExoClip预训练数据集,该数据集包含110万对同步的内向-外向视频文本对,源自Ego-Exo4D,以及从多个来源收集的指令微调数据集EgoIT,以增强模型的指令遵循能力。基于这些数据集,我们提出了一种迁移策略,并进一步设计了一个分阶段的映射学习流程,包括演示者自我准备、演示者-学习者指导和学习者自我练习三个阶段。在多样化的内向任务上进行了广泛的实验,结果显示,现有MLLM在内向视频理解上表现不足,而我们的模型显著优于这些领先模型。
英文摘要
AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric) vision, overlooking the unique challenges of first-person (egocentric) videos. Additionally, high acquisition costs limit data size, impairing MLLM performance. To address these challenges, we propose learning the mapping between exocentric and egocentric domains, leveraging the extensive exocentric knowledge within existing MLLMs to enhance egocentric video understanding. To this end, we introduce Ego-ExoClip, a pre-training dataset comprising 1.1M synchronized ego-exo clip-text pairs derived from Ego-Exo4D, together with the instruction-tuning dataset EgoIT, which is collected from multiple sources to enhance the model's instruction-following capabilities. Building upon the datasets, we propose a migration strategy and further design a progressive mapping learning pipeline with three stages: Demonstrator Self-Preparation, Demonstrator-Learner Guidance, and Learner Self-Practice. Extensive experiments across diverse egocentric tasks reveal that existing MLLMs perform inadequately in egocentric video understanding, while our model significantly outperforms these leading models.