arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.20169cs.CVcs.MM

EgoForge:面向目标的自体世界模拟器

EgoForge: Goal-Directed Egocentric World Simulator

Yifan Shen, Jiateng Liu, Xinzhuo Li, Yuanzhe Liu, Bingxuan Li, Houze Yang, Wenqi Jia, Yijiang Li, Tianjiao Yu, James Matthew Rehg, Xu Cao, Ismini Lourentzou

首次发表
浏览论文内容

中文总结 AI 辅助

EgoForge通过最小静态输入生成连贯的第一人称视频,结合VideoDiffusionNFT提升意图对齐和时间一致性,实现在动态环境模拟中的优越性能。

中文摘要 AI 辅助

生成世界模型在模拟动态环境方面显示出潜力,但单目视频仍具挑战性,因视角变化快、频繁手-物交互和目标导向过程依赖于潜在人类意图。现有方法或专注于手部中心指令合成,或进行静态视角转换而不建模动作动态,或依赖密集监督。本文提出EgoForge,一种面向目标的自体世界模拟器,从最小静态输入生成连贯的第一人称视频:单个自体图像、高层指令和可选辅助非自体视角。为提高意图对齐和时间一致性,我们提出VideoDiffusionNFT,一种轨迹级奖励引导的细化方法,优化目标完成、时间因果性、场景一致性和感知保真度。大量实验表明,EgoForge在语义对齐、几何稳定性及运动保真度上优于强基线,并在现实世界智能眼镜实验中表现出鲁棒性。

英文摘要

Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, and synchronized multi-camera capture. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal coherence, we introduce GRAFT, a trajectory-level diffusion refinement method that uses positive and negative rollout distributions, derived from goal, temporal, scene-consistency, and perceptual rewards, to steer the diffusion velocity field toward coherent, goal-complete egocentric simulations. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and performs robustly in real-world smart-glasses experiments.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • University of California San Diego(加州大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

↑