arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.17896cs.CVcs.AI

EgoWorld:利用丰富的外参照观测将外参照视角转换为自身参照视角

EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations

Junho Park, Andrew Sangwoo Ye, Taein Kwon

首次发表 更新
浏览论文内容

中文总结 AI 辅助

EgoWorld通过重建丰富的外参照观测来实现外参照到自身参照视角的转换,展示了在增强现实、虚拟现实和机器人应用中的先进性能和鲁棒性。

中文摘要 AI 辅助

自身视角视觉对于人类和机器视觉理解至关重要,特别是在捕捉所需的手-物体交互细节以完成操作任务。将第三人称观点转换为第一人称观点显著促进了增强现实(AR)、虚拟现实(VR)和机器人应用。然而,当前的外参照到自身参照转换方法受到其依赖2D线索、同步多视角设置以及不现实假设的限制,例如在推理过程中需要初始自身参照框架和相对相机姿态。为克服这些挑战,我们引入了EgoWorld,一种新的框架,能够从丰富的外参照观测中重建自身参照视角,包括点云、3D手姿态和文本描述。我们的方法从估计的外参照深度图中重建点云,将其重新投影到自身参照视角,并然后应用扩散模型生成密集且语义连贯的自身参照图像。在四个数据集(即H2O、TACO、Assembly101和Ego-Exo4D)上评估,EgoWorld实现了最先进的性能,并展示了对新对象、动作、场景和主体的稳健泛化能力。此外,EgoWorld在真实世界示例中表现出鲁棒性,突显了其实际应用价值。项目页面可在https://redorangeyellowy.github.io/EgoWorld/上获取。

英文摘要

Egocentric vision is essential for both human and machine visual understanding, particularly in capturing the detailed hand-object interactions needed for manipulation tasks. Translating third-person views into first-person views significantly benefits augmented reality (AR), virtual reality (VR) and robotics applications. However, current exocentric-to-egocentric translation methods are limited by their dependence on 2D cues, synchronized multi-view settings, and unrealistic assumptions such as the necessity of an initial egocentric frame and relative camera poses during inference. To overcome these challenges, we introduce EgoWorld, a novel framework that reconstructs an egocentric view from rich exocentric observations, including point clouds, 3D hand poses, and textual descriptions. Our approach reconstructs a point cloud from estimated exocentric depth maps, reprojects it into the egocentric perspective, and then applies diffusion model to produce dense, semantically coherent egocentric images. Evaluated on four datasets (i.e., H2O, TACO, Assembly101, and Ego-Exo4D), EgoWorld achieves state-of-the-art performance and demonstrates robust generalization to new objects, actions, scenes, and subjects. Moreover, EgoWorld exhibits robustness on in-the-wild examples, underscoring its practical applicability. Project page is available at https://redorangeyellowy.github.io/EgoWorld/.

发表机构

  • AI Lab, LG Electronics(LG电子人工智能实验室)
  • KAIST(韩国科学技术院)
  • Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑