发表机构
EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对潜在SDS生成3D时的像素漂移问题,提出轻量级VAE一致性梯度修复方法PixSDS,减少结构化伪影并保留语义内容。
AI 中文摘要
分数蒸馏采样(Score Distillation Sampling,SDS)通过预训练扩散先验优化渲染图像,实现文本到3D生成,但潜在SDS常产生结构化色彩伪影和高频纹理噪声。我们识别出潜在SDS的一种失效模式,由VAE诱导的像素漂移导致:优化后的图像可沿VAE编码器约束较弱的像素空间方向移动,其潜在表示仍保持干净且语义有意义,而图像本身却累积可见伪影。我们通过受控2D SDS实验、仅VAE优化实验,以及一项简化分析支持该诊断,该分析表明当映射回像素的逆映射约束不足时,类编码器的潜在目标会放大图像空间噪声。基于该观察,我们提出PixSDS,一种轻量级VAE一致性梯度修复方法。PixSDS解码潜在SDS的前瞻步,并将解码后的图像作为像素空间优化的干净方向,减少VAE不一致方向上的运动,且无需重新训练扩散模型、更改渲染器或替换SDS目标。2D优化和文本到3D生成实验表明,PixSDS在保留语义内容的同时大幅减少了结构化伪影。代码可在此httpsURL获取。
英文摘要
Score Distillation Sampling (SDS) enables text-to-3D generation by optimizing rendered images with a pretrained diffusion prior, but latent SDS often produces structured color artifacts and high-frequency texture noise. We identify a failure mode of latent SDS caused by VAE-induced pixel drift: the optimized image can move along pixel-space directions that are weakly constrained by the VAE encoder, so its latent representation remains clean and semantically meaningful while the image itself accumulates visible artifacts. We support this diagnosis with controlled 2D SDS experiments, VAE-only optimization, and a simplified analysis showing that encoder-like latent objectives can amplify image-space noise when the inverse mapping to pixels is underconstrained. Motivated by this observation, we propose PixSDS, a lightweight VAE-consistent gradient repair method. PixSDS decodes a latent SDS lookahead step and uses the decoded image as a clean direction for pixel-space optimization, reducing motion in VAE-inconsistent directions without retraining the diffusion model, changing the renderer, or replacing the SDS objective. Experiments in 2D optimization and text-to-3D generation show that PixSDS substantially reduces structured artifacts while preserving semantic content. Code is publicly available at https://sevashasla.github.io/pixsds-webpage/.