发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Luce,一种用于3D资产生成的可重光照高斯模型,通过多模态高斯云结合PBR材质,在Toys4K数据集及自研AI生成图像基准上均实现了优于基线的单图像到3D生成性能,可保留精细细节。
AI 中文摘要
高保真图像到3D生成需要同时捕捉几何与外观的3D表示,为支持重光照及集成到标准渲染管线,该表示应包含基于物理的渲染(PBR)模态,如反照率(albedo)、金属粗糙度(metallic-roughness)和表面法线。本文提出Luce,一种在体素化多模态高斯云中统一几何与PBR材质的3D表示,每个模态使用专用高斯基元;变分自编码器将该表示压缩为统一的材质感知潜在空间,整流流Transformer基于预训练图像编码器的多层特征生成该潜在空间,这些特征保留语义上下文与精细空间细节;潜在空间随后解码为可重光照的PBR高斯模型,以及带切线空间法线贴图的可选纹理网格。在Toys4K数据集上,Luce实现了最先进的单图像到3D生成,相比最强基线的FID降低28%;本文还引入了AI生成图像基准,在该基准上Luce的CLIP图像对齐分数优于最佳基线(0.8519对比0.8299),生成的可重光照资产几何准确、材质忠实,能保留文字、标志和铭文等精细细节。
英文摘要
High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. However, preserving fine detail across the physically based rendering (PBR) modalities needed for relighting remains challenging. To address this, we propose Luce, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for albedo, metallic-roughness, and surface normals. A variational autoencoder compresses this representation into a unified material-aware latent space. A rectified-flow transformer generates this latent from a single image using multi-layer features from a pretrained image encoder that preserve both semantic context and fine spatial detail. The latent is then decoded into relightable PBR Gaussians and an optional textured mesh with a tangent-space normal map. On Toys4K, Luce achieves state-of-the-art single-image-to-3D generation, improving FID by 28% over the strongest baseline. We further evaluate Luce on a benchmark of AI-generated images depicting diverse subjects and materials, where it improves the CLIP image-alignment score over the best baseline (0.8519 vs. 0.8299). Luce generates relightable, geometrically accurate, and materially faithful assets that preserve fine details such as text, logos, and inscriptions.
Comments28 pages, 19 figures, 5 tables