发表机构
School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本生成360度动态人类的现有方法存在成本高、一致性差的问题,本文提出4DHumanDiff扩散框架,直接基于文本生成4DGS表示的动态人类,结合新数据集与技术,实现高效且一致性更好的生成。
AI 中文摘要
从文本提示生成高质量的360度动态人类资产颇具挑战性。现有方法通常先合成单目或多视角视频,再拟合4D表示,这种方式成本高昂,且常导致几何结构不完整或视角不一致的渲染结果。我们提出4DHumanDiff,这是一种扩散框架,可直接基于文本提示生成以4D高斯溅射(4DGS)表示的动态人类。通过端到端建模结构化的4D表示空间,4DHumanDiff避免了视频预生成和逐场景重建,更适合生成视角一致且时间连贯的资产。该模型采用带时间注意力的3D U-Net骨干网络,用于感知运动的生成。我们还构建了包含60000个高质量文本-4DGS对的大规模数据集,并引入2D正则化和无需训练的4D插值,以提升渲染质量和运动平滑度。实验表明,4DHumanDiff可在一分钟内生成一致的360度动态人类,实现更优的时间和多视角一致性,且推理时间缩短10倍以上。
英文摘要
Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and often causes incomplete geometry or view-inconsistent renderings. We present 4DHumanDiff, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting (4DGS) from text prompts. By modeling the structured 4D representation space end-to-end, 4DHumanDiff avoids video pre-generation and per-scene reconstruction, making it better suited for view-consistent and temporally coherent asset generation. The model uses a 3D U-Net backbone with temporal attention for motion-aware generation. We further construct a large-scale text-to-4DGS dataset with 60,000 high-quality pairs, and introduce 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness. Experiments show that 4DHumanDiff generates consistent 360-degree dynamic humans within one minute, achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.
Comments14 pages