arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38615cs.CVcs.AI

Exo2EgoHOI:手-物交互感知的外视到自我中心视频生成

Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation

Hongjia Zhai, Xiyu Zhang, Haoran Zhang, Zhichao Ye, Haomin Liu, Guofeng Zhang, Ian Reid, Xingxing Zuo

首次发表
浏览论文内容

中文总结 AI 辅助

提出Exo2EgoHOI框架,利用4D HOI先验和分解门控交叉注意力,在外视到自我中心视频生成中保留手-物交互和目标一致性,在ARCTIC-HOI和Ego-Exo4D上显著提升性能。

中文摘要 AI 辅助

人类操作的第一人称视频为具身智能提供了宝贵的视觉经验,然而大规模采集此类数据成本高昂。外视到自我中心视频生成通过将丰富的第三人称操作视频转化为第一人称观察,提供了一种可扩展的替代方案。然而,现有方法由于缺乏细粒度的交互引导和弱的目标中心锚定,往往难以在大的视角变化下忠实保留演示的手-物交互(HOI)。我们提出了Exo2EgoHOI,一个HOI感知的视频生成框架,用于保留交互的外视到自我中心转换。为保留细粒度HOI,我们引入了一个统一的4D HOI先验,结合了场景几何、关节手部渲染和密集的手-物关系场,并配有一个双分支残差适配器,将结构和关系线索注入视频生成主干。为保持目标一致性,我们引入了分解门控交叉注意力,分别编码目标和背景参考,并自适应地整合全局语义和局部外观特征作为目标中心锚点。在ARCTIC-HOI和Ego-Exo4D上的实验表明,在保持竞争性视觉保真度的同时,目标一致性和HOI保留方面有显著改进。特别是在ARCTIC-HOI上,相对于各自的最佳基线结果,Exo2EgoHOI将目标mIoU提高了32.3%,并将MPJPE和PA-MPJPE分别降低了34.7%和50.0%。项目页面:此https URL。

英文摘要

Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results. Project page: https://rcl-robotics.github.io/Exo2EgoHOI/.

发表机构

  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • Zhejiang University(浙江大学)
  • InSpatio

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑