arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23410cs.CVcs.LG

基于Next-Scale Transformer的人脸真实感新视角合成

Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

  • École polytechnique fédérale de Lausanne(洛桑联邦理工学院)
  • Meta

机构由 AI 辅助整理,请以论文原文为准。

Federico Stella, Fei Jiang, Zhongshi Jiang, Zohar Barzelay, Emanuel Garbin, Amin Jourabloo, Liuhao Ge

AI总结:

本研究将next-scale自回归范式适配人脸新视角合成,结合Transformer模型实现高保真多视角人脸3D模型生成,较扩散模型需更少专用训练数据,提升了感知与跨视角一致性。

AI中文摘要:

人脸的真实感新视角合成在高空间分辨率及多目标相机场景下仍具挑战性,此时保留身份、精细外观细节与几何一致性至关重要。我们基于next-scale自回归范式,将其适配于人中心视角合成任务,使其支持更高图像分辨率、多视角输出及单次前向传播中更强的跨视角一致性。我们在涵盖多样身份与服饰的合成人脸数据集上进行训练。与扩散模型不同,该范式无需2D预训练,且得益于其next-scale架构,可从低分辨率通用预训练中获益,仅在最后训练阶段使用全尺寸专用图像。这使我们的架构能以更少的专用训练数据收敛,因此可使用规模更小但更贴近真实的训练数据集。所得模型可生成清晰真实的视角,还可选择同时合成多个新视点以提升跨视角一致性。经实验,我们在人脸主体上观察到感知保真度与跨视角一致性的提升,证明next-scale自回归是可扩展多输出人视角合成的有效骨干。我们还将该流程与现有基于Transformer的模型结合,用于从多视角人脸输入中进行像素对齐的3D高斯提升,从而生成准确且真实感强的人脸3D模型。

英文摘要:

Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving identity, fine appearance details, and geometric coherence is critical. We build on the next-scale autoregressive paradigm and adapt it for human-centric view synthesis by enabling higher image resolutions, multi-view outputs and stronger cross-view consistency in a single forward pass. We train on a synthetic dataset of human faces spanning diverse identities and apparel. Contrary to diffusion models, this paradigm does not need 2D pre-training and, thanks to its next-scale architecture, it benefits from lower-resolution, general-purpose pre-trainings, with the full-sized purpose-specific images being used only in the last training stages. This enables our architecture to converge with a smaller amount of purpose-specific training data, allowing us to use a smaller but more realistic training dataset. The resulting model produces sharp and realistic views, with the option to synthesize multiple novel viewpoints simultaneously for improved agreement across views. Empirically, we observe gains in perceptual fidelity and cross-view coherence on human subjects, demonstrating that next-scale autoregression is an effective backbone for scalable, multi-output human view synthesis. We also couple our pipeline with an existing transformer-based model for pixel-aligned 3D gaussian lifting from multi-view facial inputs, resulting in accurate and photorealistic 3D models of human faces.

↑