发表机构
Huazhong University of Science and Technology; Wuhan Institute of Technology; Xiaomi EV(华中科技大学; 武汉工程大学; 小米汽车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究虚拟试穿问题,提出STAR-VTON框架,通过解耦潜在空间结构合成与像素空间细节恢复,结合匹配信息细化器,实现高效高保真虚拟试穿,其中VAR-VTON效率高,像素空间细化器能有效恢复细节。
AI 中文摘要
虚拟试穿(VTON)是一个双条件图像生成问题,需要准确保留人物并忠实进行服装变形和细节合成。基于扩散的VTON方法在压缩潜在空间中联合建模这些因素,但因固有潜在压缩存在高频细节损失。近期视觉自回归(VAR)模型为高质量生成提供了有前景的替代方案,但因缺乏有效双条件机制未用于VTON。为此引入VAR-VTON,虽有效但在保留细粒度服装细节上仍有困难。进而提出STAR-VTON,通过解耦潜在空间结构合成与像素空间细节恢复,借助匹配信息细化器建立对应关系来恢复细节。实验表明STAR-VTON实现了效率与保真度的良好权衡,VAR-VTON比基于扩散的方法快至少4倍且不降低质量,像素空间细化器能有效恢复细节并可作为即插即用模块惠及现有VTON方法。
英文摘要
Virtual try-on (VTON) is a bi-conditional image generation problem that requires not only accurate person preservation but also faithful garment deformation and detail synthesis. Diffusion-based VTON methods can jointly model these factors in a compressed latent space, but suffer from high-frequency detail loss due to inherent latent compression, even with costly multi-step denoising. Recent visual autoregressive (VAR) models offer a promising alternative for high-quality generation with faster inference, yet remain unexplored for VTON due to the lack of effective bi-conditioning mechanisms. To bridge this gap, we first introduce VAR-VTON, a VAR-based VTON model that incorporates garment conditioning and structural guidance for efficient latent-space VTON. Despite its efficacy, latent-space generation still struggles to preserve fine-grained garment details. We argue that different VTON sub-tasks should be addressed in different representation spaces: structural synthesis such as garment warping and person layout is suited to the latent space, whereas fine-grained detail recovery should be tackled in the pixel space. Motivated by this insight, we further propose STAR-VTON, a Two-Stage AutoRegressive framework that builds upon VAR-VTON by decoupling latent-space structural synthesis from pixel-space detail recovery. Our idea is to resort to a matching-informed refiner to establish dense correspondences between the stage-one generation and the source garment to directly map fine-grained pixel-space details. Extensive experiments show that STAR-VTON achieves an impressive efficiency-fidelity trade-off: VAR-VTON runs at least $4\times$ faster than diffusion-based counterparts without degrading quality, and the pixel-space refiner effectively restores fine details and acts as a plug-and-play module that can benefit existing VTON approaches.