arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Ego-Forge:无文本与几何注意力的外部到自我中心视频生成

Ego-Forge: Text and Geometric-Attention Free Exo-to-Egocentric Video Generation

Mohammad Mahdi, Luc Van Gool, Danda Pani Paudel

arXiv 2609.35368首次发表:更新:

发表机构

INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学圣克莱门特奥赫里德分校INSAIT)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Ego-Forge提出无文本和几何注意力的外部到自我中心视频生成框架,通过动态标题生成和扩大训练规模,在Ego-Exo4D上达到最先进性能,无需外部标注且泛化到野外场景。

AI 中文摘要

外部到自我中心视频生成旨在根据第三人称视频和目标头部轨迹,合成一个人从自身视角看到的内容。该任务需要在大的视角变化间迁移外观和语义,同时幻觉出外部相机从未观察到的内容。现有方法要么施加额外的输入要求,如真实初始自我中心帧或多个同步外部视角,要么局限于类别特定的设置。EgoX是首个解决跨活动和野外泛化的方法,但在推理时需要人类提供不存在的自我中心视图的标题,并引入了计算昂贵的几何引导注意力偏差,该偏差可能传播重建误差并抑制文本和视觉上下文。因此,我们提出Ego-Forge,一个无标题和无偏差的外部到自我中心生成框架。它引入了动态标题生成,直接从模型隐藏状态中派生条件令牌,并使其适应扩散时间步和网络深度,取代外部文本条件。通过将训练规模扩大一个数量级并使用所有可用的外部视角,Ego-Forge隐式学习跨视角对应关系,消除了对几何引导注意力的需求,仅需轻量深度先验。Ego-Forge在Ego-Exo4D上达到最先进性能,端到端运行更快,推理时无需外部标注,并泛化到野外场景,包括过度依赖几何会阻碍外观推断的情况。我们的模型和源代码将公开发布。

英文摘要

Exo-to-egocentric video generation aims to synthesize what a person sees from their own viewpoint given third-person footage and a target head trajectory. The task requires transferring appearance and semantics across large viewpoint changes while hallucinating content never observed by the exocentric camera. Existing approaches either impose additional input requirements, such as a ground-truth initial egocentric frame or multiple synchronized exocentric views, or remain limited to category-specific settings. EgoX is the first to address cross-activity and in-the-wild generalization, but requires a human-provided caption of the non-existent egocentric view at inference and introduces a computationally expensive geometry-guided attention bias that can propagate reconstruction errors and suppress textual and visual context. We therefore propose \textbf{Ego-Forge}, a caption-free and bias-free framework for exo-to-egocentric generation. It introduces \textit{Dynamic Captioning}, which derives conditioning tokens directly from the model's hidden states and adapts them to the diffusion timestep and network depth, replacing external text conditioning. By scaling training by an order of magnitude and using all available exocentric viewpoints, Ego-Forge learns cross-view correspondence implicitly and eliminates the need for geometry-guided attention, requiring only a lightweight depth prior. Ego-Forge achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearance inference. Our model and source code will be made publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑