发表机构
National Yang Ming Chiao Tung University; Alaya Lab(国立阳明交通大学; Alaya实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OuroWorld是一种无遮罩框架,通过提出Inconsistency-Robust Periodic 4DGS方法,将静态3D高斯溅射场景转化为可从任意视点无缝循环的3D动态影像,在39个场景的用户研究中胜率达70.8%-99.0%,性能优于基线方法。
AI 中文摘要
现有3D世界模型可生成逼真、可探索的场景,但这些场景始终处于时间冻结状态。OuroWorld是一种无遮罩框架,能将任意静态3D高斯溅射(3D Gaussian Splatting)场景转化为3D动态影像,即从任意视点出发都能无缝循环、带有生动多样运动的动态场景。视觉语言模型会推断合理的动态信息,并引导视频模型合成参考视频,我们将该参考视频提升并补全为多视图视频。为从这种不完善的监督中学习,我们提出了不一致鲁棒周期性4D高斯溅射(Inconsistency-Robust Periodic 4DGS):傅里叶级数变形场通过构造保证循环性,而锚定在参考视图的接地漂移场(Grounded Drift Field)则吸收跨视图不一致性。与此前仅能处理类流体运动的欧拉方法不同,我们能捕捉一般变形、物体运动和光照变化。我们引入了无真值评估,涵盖生动性、自然度、循环接缝一致性和场景质量。在39个重建和生成的场景上,OuroWorld的表现优于所有基线方法,在用户研究比较中赢得了70.8%-99.0%的胜率。项目页面:this https URL
英文摘要
Recent 3D world models generate photorealistic, explorable scenes that remain frozen in time. OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint. A vision-language model infers plausible dynamics and guides a video model to synthesize a reference video, which we lift and complete into multi-view videos. To learn from this imperfect supervision, we propose Inconsistency-Robust Periodic 4DGS: a Fourier-series deformation field guarantees looping by construction, while a Grounded Drift Field anchored at the reference view absorbs cross-view inconsistency. Unlike prior Eulerian methods limited to fluid-like motion, we capture general deformation, object motion, and illumination change. We introduce a ground-truth-free evaluation covering vividness, naturalness, loop seam coherence, and scene quality. On 39 reconstructed and generated scenes, OuroWorld outperforms all baselines and wins 70.8%-99.0% of user-study comparisons. Project page: https://ouroworld.userwei.com
CommentsProject page: https://ouroworld.userwei.com