arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11722cs.CV

重新审视“化身即图像”:高保真配准才是关键

Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need

Margaret Kostyrko, Yuxuan Xue, Garvita Tiwari, Gerard Pons-Moll

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AvaImg多阶段优化流程,通过符号环绕数约束和三级效率级联实现高保真SMPL-X+D配准,在六个数据集上优于基线,并验证了其UV表示与图像基础模型的兼容性。

中文摘要 AI 辅助

将三维着装人体表示为标准化二维UV纹理和基于底层人体模型的位移图,这一方法长期以来一直受到研究。这种紧凑的表示方法颇具吸引力,因为它使得预训练的图像网络能够处理、生成和编辑三维化身,但只有当扫描数据通过高保真配准被精确对齐并建立对应关系时,这种方法才有用。这一前提条件从未被满足,我们认为这解释了以往基于UV的着装人体方法质量有限的原因。尽管其重要性不言而喻,但目前尚无公开方法能从任意着装扫描中生成带有UV纹理的高保真SMPL(-X)+D配准结果。我们提出了AvaImg,一个多阶段优化流程,以填补这一空白:它通过符号环绕数强制执行“身体在衣服内部”的约束,并通过三级效率级联(运行时间减少约10倍,存储节省约95%)使其可行,同时利用从粗到细的位移优化恢复精细的表面细节。AvaImg在六个数据集上的身体拟合、形状估计和表面配准方面均优于所有基线,生成的纹理配准结果与扫描几乎无法区分(PSNR=34.48dB)。为了验证AvaImg的“化身即图像”表示与图像基础模型的即时兼容性,我们通过冻结的FLUX VAE对我们的UV图进行自动编码。这仅增加了相对于扫描的0.76mm Chamfer误差,并表明生成的图位于自然图像分布之内,支持将二维生成先验用于三维化身生成。代码、数据和Singularity容器将在此https URL提供。

英文摘要

The representation of 3D clothed humans as standardized 2D UV texture and displacement maps over an underlying body model has long been studied. This compact representation is enticing as it enables pretrained image networks to process, generate, and edit 3D avatars, but is only useful if scans are accurately aligned and brought into correspondence via high-fidelity registration. This prerequisite has never been met, which we argue explains the limited quality of prior UV-based methods for clothed humans. Despite its significance, no public method produces high-fidelity SMPL(-X)+D registrations with UV texture from arbitrary clothed scans. We present AvaImg, a multi-stage optimization pipeline, to close this gap: it enforces body-inside-clothing constraint via signed winding numbers, made viable by a three-level efficiency cascade (~10x runtime reduced, ~95% storage saved), and recovers fine surface detail using coarse-to-fine displacement optimization. AvaImg outperforms all baselines in body fitting, shape estimation, and surface registration across six datasets, yielding textured registrations near-indistinguishable from scans (PSNR=34.48dB). For validation of AvaImg's Avatar-as-Image representation as imminently compatible with image foundation models, we auto-encode our UV maps via the frozen FLUX VAE. This achieves only 0.76mm added Chamfer error relative to scan and shows that the resulting maps lie within natural-image distributions, supporting the use of 2D generative priors for 3D avatar generation. Code, data, and Singularity containers will be at https://yuxuan-xue.com/avaimg.

发表机构

  • Tübingen AI Center(图宾根人工智能中心)
  • Max Planck Institute for Informatics(马克斯·普朗克信息学研究所)
  • University of Tübingen(图宾根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑