发表机构
CLO Virtual Fashion Inc.(CLO虚拟时尚公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DiffGI针对现有3D生成模型问题,提出端到端3D到2D映射框架,用连续函数取代二进制图,引入可微算法,训练相关模型,在薄壳3D生成中实现高保真、高精度,减少计算资源需求。
AI 中文摘要
现有3D生成模型主要依赖隐式体表示,难以处理薄壳和非流形几何。基于几何图像的方法存在分辨率依赖边界编码问题。为此提出DiffGI,它是端到端3D到2D映射框架,用连续2D截断符号距离函数取代二进制图,引入可微Marching Squares算法,训练DiffGI-VAE并实例化基于变压器的潜在扩散模型。实验表明该方法在重建保真度和边界精度上优于现有方法,且计算资源需求少。
英文摘要
Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce watertight topology and struggle to represent thin-shell and non-manifold geometries such as garments. Geometry image-based approaches offer a surface-centric alternative, but existing methods rely on discrete binary occupancy maps whose resolution-dependent boundary encoding causes staircase artifacts and information loss upon downsampling, while surface reconstruction remains a non-differentiable post-processing step disconnected from the learning pipeline. To address this, we propose Differentiable Geometry Image (DiffGI), an end-to-end 3D-to-2D mapping framework that seamlessly integrates surface representation and geometric optimization. DiffGI replaces binary maps with a continuous 2D Truncated Signed Distance Function (TSDF), which encodes boundary position at subpixel precision within a fixed grid resolution, eliminating resolution-dependent staircase artifacts even under aggressive downsampling. Building on this continuous field, we introduce a differentiable Marching Squares algorithm based on analytical linear interpolation, allowing gradients from 3D surface losses to propagate back to the 2D latent space. Leveraging this differentiable pipeline, we train a DiffGI-VAE augmented with a geometry-aware normal rendering loss to compress complex 3D surfaces into an ultra-compact 32X32 latent space, and instantiate a transformer-based latent diffusion model with a flow-matching objective on top of this space for conditional 3D generation. Extensive experiments on garment and object datasets demonstrate that our method achieves superior reconstruction fidelity and boundary precision compared to prior geometry-image and voxel-based approaches, while requiring significantly fewer computational resources.
CommentsAccepted to ECCV 2026