EgoForge:面向目标的自体世界模拟器
EgoForge: Goal-Directed Egocentric World Simulator
浏览论文内容
中文总结 AI 辅助
EgoForge通过最小静态输入生成连贯的第一人称视频,结合VideoDiffusionNFT提升意图对齐和时间一致性,实现在动态环境模拟中的优越性能。
中文摘要 AI 辅助
生成世界模型在模拟动态环境方面显示出潜力,但单目视频仍具挑战性,因视角变化快、频繁手-物交互和目标导向过程依赖于潜在人类意图。现有方法或专注于手部中心指令合成,或进行静态视角转换而不建模动作动态,或依赖密集监督。本文提出EgoForge,一种面向目标的自体世界模拟器,从最小静态输入生成连贯的第一人称视频:单个自体图像、高层指令和可选辅助非自体视角。为提高意图对齐和时间一致性,我们提出VideoDiffusionNFT,一种轨迹级奖励引导的细化方法,优化目标完成、时间因果性、场景一致性和感知保真度。大量实验表明,EgoForge在语义对齐、几何稳定性及运动保真度上优于强基线,并在现实世界智能眼镜实验中表现出鲁棒性。
英文摘要
Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, and synchronized multi-camera capture. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal coherence, we introduce GRAFT, a trajectory-level diffusion refinement method that uses positive and negative rollout distributions, derived from goal, temporal, scene-consistency, and perceptual rewards, to steer the diffusion velocity field toward coherent, goal-complete egocentric simulations. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and performs robustly in real-world smart-glasses experiments.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- University of California San Diego(加州大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。