arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

4DAnyone:从随意的单目视频中创建4D虚拟人物

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

Yudong Jin, Tao Xie, Qihang Zhang, Zehong Shen, Zhen Xu, Yujun Shen, Hujun Bao, Xiaowei Zhou, Yinghao Xu

arXiv 2608.20335首次发表:更新:

发表机构

State Key Lab of CAD&CG, Zhejiang University; Robbyant; Chinese University of Hong Kong; Ant Group; Hong Kong University of Science and Technology(浙江大学CAD&CG国家重点实验室; Robbyant; 香港中文大学; 蚂蚁集团; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

4DAnyone是一种从单目视频重建4D人体的框架,通过RCP和TCR解决多视图一致性问题,在相关数据集上优于现有方法且泛化性好。

AI 中文摘要

我们提出了4DAnyone,这是一种通过生成重建级多视图一致视频并将其提升为4D高斯溅射(4DGS),从未校准的单目视频中重建4D人体的框架。现有的相机控制视频扩散模型可合成合理的新视图视频,但在扩展到4DGS重建所需的数十个目标视图时无法保持一致性。我们将此失败识别为有限注意力上下文问题:当目标视图超过单个DiT前向传递的容量时,必须将其拆分为组,这会暴露两个耦合瓶颈。在参考上下文方面,对所有先前生成的视图进行条件化的复杂度为O(N),削弱了跨视图外观引导。在目标上下文方面,不相交的组无法直接交换信息,导致全局结构漂移。4DAnyone通过两种互补设计解决了这两个瓶颈:参考上下文打包(RCP)将不断增长的参考视图压缩为固定长度的混合分辨率上下文,具有O(1)的参考上下文复杂度;目标上下文路由(TCR)在去噪过程中旋转目标视图分组,以便在高噪声步骤中跨组共享上下文,并在低噪声步骤中稳定细节。我们还使用内部游戏引擎构建了MVGameHuman数据集,并将其与灯光舞台和野外视频数据集结合用于训练。在DNA-Rendering和DyMVHumans上的实验表明,4DAnyone在新视图视频质量和下游4DGS重建方面均优于现有方法,具有稳健的野外泛化能力。请访问我们的项目页面查看视频结果和源代码:this https URL

英文摘要

We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existing camera-controlled video diffusion models synthesize plausible novel-view videos but fail to maintain consistency when scaled to the tens of target views required for 4DGS reconstruction. We identify this failure as a bounded-attention-context problem: when target views exceed the capacity of a single DiT forward pass, they must be split into groups, exposing two coupled bottlenecks. On the reference-context side, conditioning on all previously generated views grows as $O(N)$, weakening cross-view appearance guidance. On the target-context side, disjoint groups cannot directly exchange information, causing global structural drift. 4DAnyone addresses both bottlenecks with two complementary designs: Reference Context Packing (RCP) compresses growing reference views into a fixed-length mixed-resolution context with $O(1)$ reference-context complexity, while Target Context Routing (TCR) rotates target-view groupings during denoising to share context across groups at high-noise steps and stabilize details at low-noise steps. We further build the MVGameHuman dataset using our in-house game engine and combine it with light-stage and in-the-wild video datasets for training. Experiments on DNA-Rendering and DyMVHumans show that 4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization. See our project page for video results and source code: https://4danyone.github.io.

CommentsProject page: https://4danyone.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑