arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2403.09194cs.CV

意图驱动的自我中心到外部中心视频生成

Intention-driven Ego-to-Exo Video Generation

  • USTC(中国科学技术大学)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Hongchen Luo, Kai Zhu, Wei Zhai, Yang Cao

更新

AI总结:

针对视角剧变下现有视频生成方法失效的问题,提出意图驱动框架IDE,利用动作意图作为视角无关表示,结合跨视角特征感知与动作描述,指导扩散模型生成外部中心视频,实验验证其优于现有模型。

AI中文摘要:

自我中心到外部中心视频生成是指根据自我中心视频生成相应的外部中心视频,在增强现实/虚拟现实和具身智能中具有重要应用价值。得益于扩散模型技术的进步,视频生成领域已取得显著进展。然而,现有方法建立在相邻帧之间的时空一致性假设之上,而在自我中心到外部中心场景中,由于视角的剧烈变化,该假设无法满足。为此,本文提出了一种意图驱动的自我中心到外部中心视频生成框架(IDE),利用由人体运动和动作描述组成的动作意图作为视角无关的表示来指导视频生成,从而保持内容和运动的一致性。具体而言,首先通过多视角立体匹配估计自我中心头部轨迹。然后,引入跨视角特征感知模块以建立外部中心视角与自我中心视角之间的对应关系,指导轨迹变换模块从头部轨迹推断人体全身运动。同时,我们提出一个动作描述单元,将动作语义映射到与外部中心图像一致的特征空间中。最后,推断出的人体运动和高级动作描述共同指导扩散模型反向过程中外部中心运动和交互内容(即对应的光流和遮挡图)的生成,最终将其扭曲为相应的外部中心视频。我们在包含多样化外部-自我中心视频对的相关数据集上进行了广泛实验,我们的IDE在主观和客观评估中均优于最先进的模型,证明了其在自我中心到外部中心视频生成中的有效性。

英文摘要:

Ego-to-exo video generation refers to generating the corresponding exocentric video according to the egocentric video, providing valuable applications in AR/VR and embodied AI. Benefiting from advancements in diffusion model techniques, notable progress has been achieved in video generation. However, existing methods build upon the spatiotemporal consistency assumptions between adjacent frames, which cannot be satisfied in the ego-to-exo scenarios due to drastic changes in views. To this end, this paper proposes an Intention-Driven Ego-to-exo video generation framework (IDE) that leverages action intention consisting of human movement and action description as view-independent representation to guide video generation, preserving the consistency of content and motion. Specifically, the egocentric head trajectory is first estimated through multi-view stereo matching. Then, cross-view feature perception module is introduced to establish correspondences between exo- and ego- views, guiding the trajectory transformation module to infer human full-body movement from the head trajectory. Meanwhile, we present an action description unit that maps the action semantics into the feature space consistent with the exocentric image. Finally, the inferred human movement and high-level action descriptions jointly guide the generation of exocentric motion and interaction content (i.e., corresponding optical flow and occlusion maps) in the backward process of the diffusion model, ultimately warping them into the corresponding exocentric video. We conduct extensive experiments on the relevant dataset with diverse exo-ego video pairs, and our IDE outperforms state-of-the-art models in both subjective and objective assessments, demonstrating its efficacy in ego-to-exo video generation.

↑