发表机构
Seoul National University(首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现实场景中低重叠度相机的4D人体场景重建问题,提出StudioRecon方法,通过解耦背景和人体,利用视频扩散模型强化背景监督,结合跨视图关联等初始化人体,经递归增强模块避免伪影,实现了高水准新视图合成及相关应用。
AI 中文摘要
现有的动态人体动作容积捕获通过密集相机阵列实现高保真度。但在现实场景中,只有少数低重叠度相机,会降低输出质量并留下大面积未观测区域。近期的4D重建方法聚焦低重叠度设置,但在观测不足区域仍有明显伪影。视频扩散模型也存在几何不一致问题。为此提出StudioRecon,通过解耦背景和人体从稀疏、低重叠度相机重建4D人体场景。利用视频扩散模型合成数百个相机控制的新视图强化背景监督,通过跨视图身份关联和三角测量多视图关键点拟合稳健初始化可变形高斯人体,递归增强模块避免残留伪影。在四个真实世界数据集上实现了最新的新视图合成,并展示了新轨迹渲染和人体替换等应用。
英文摘要
Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multi-view keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art novel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement.
CommentsAccepted to SIGGRAPH Conference Papers '26. First two authors contributed equally. Project page: https://sisyphm.github.io/studiorecon-page/