连接视角:关于自我中心-异我中心视觉的跨视角协作智能综述
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
- State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)
- University of Tokyo(东京大学)
- Zhejiang University(浙江大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文综述自我中心与异我中心视觉的跨视角协作智能,梳理三类研究方向、任务、数据集及局限,并展望未来发展。
AI中文摘要:
从自我中心(第一人称)和异我中心(第三人称)两种视角感知世界是人类认知的基础,能够实现对动态环境丰富且互补的理解。近年来,让机器利用这双重视角的协同潜力已成为视频理解中一个引人注目的研究方向。在本综述中,我们全面回顾了从异我中心和自我中心两种视角开展的视频理解研究。我们首先强调整合自我中心与异我中心技术的实际应用,展望它们在不同领域中的潜在协作。随后,我们确定实现这些应用的关键研究任务。接下来,我们系统梳理并回顾近期进展,将其组织为三个主要研究方向:(1)利用自我中心数据增强异我中心理解;(2)利用异我中心数据改进自我中心分析;(3)统一两种视角的联合学习框架。对于每个方向,我们分析了多样化的任务和相关工作。此外,我们讨论了支持双视角研究的基准数据集,评估其范围、多样性和适用性。最后,我们讨论当前工作的局限性,并提出有前景的未来研究方向。通过综合两种视角的洞见,我们的目标是激发视频理解和人工智能的进展,使机器更接近以类人方式感知世界。相关工作的 GitHub 仓库可在 https://github.com/ayiyayi/Awesome-Egocentric-and-Exocentric-Vision 找到。
英文摘要:
Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the machines to leverage the synergistic potential of these dual perspectives has emerged as a compelling research direction in video understanding. In this survey, we provide a comprehensive review of video understanding from both exocentric and egocentric viewpoints. We begin by highlighting the practical applications of integrating egocentric and exocentric techniques, envisioning their potential collaboration across domains. We then identify key research tasks to realize these applications. Next, we systematically organize and review recent advancements into three main research directions: (1) leveraging egocentric data to enhance exocentric understanding, (2) utilizing exocentric data to improve egocentric analysis, and (3) joint learning frameworks that unify both perspectives. For each direction, we analyze a diverse set of tasks and relevant works. Additionally, we discuss benchmark datasets that support research in both perspectives, evaluating their scope, diversity, and applicability. Finally, we discuss limitations in current works and propose promising future research directions. By synthesizing insights from both perspectives, our goal is to inspire advancements in video understanding and artificial intelligence, bringing machines closer to perceiving the world in a human-like manner. A GitHub repo of related works can be found at https://github.com/ayiyayi/Awesome-Egocentric-and-Exocentric-Vision.