超越全局真实感:基于维度化服装保真度评估的虚拟试穿评估与优化
Beyond Global Realism: Virtual Try-On Evaluation and Optimization with Dimension-wise Garment Fidelity Assessment
浏览论文内容
中文总结 AI 辅助
针对现有虚拟试穿评估指标难以捕捉服装多维保真度的问题,提出DAT框架,分解服装一致性为7个维度并训练专用评估模型,其性能优于多款专有模型,还可用于优化Qwen-Image-Edit的虚拟试穿生成。
中文摘要 AI 辅助
虚拟试穿(Virtual Try-On, VTON)不仅需要生成效果逼真,还需忠实保留服装特征。然而,现有评估指标如PSNR、SSIM、KID和FID难以衡量生成服装与参考服装的一致性,尤其无法捕捉服装保真度的多维特征。为解决这一问题,我们提出DAT:虚拟试穿的维度化评估框架,该框架将服装一致性分解为7个可解释维度:轮廓、颜色、领口与袖型、主要装饰与结构、材质纹理、细节保真度及logo保留,每个维度均被建模为离散属性级预测任务。为训练该专用评估模型,我们采用两阶段学习范式:先对50K样本进行大规模弱监督训练,再通过多模型投票获取10K更高质量标注进行微调。此外,我们使用加权交叉熵损失缓解各评估维度间严重的标签不平衡问题。除作为评估框架外,该评估模型可集成至Qwen-Image-Edit的虚拟试穿强化学习优化中,训练过程中会自适应聚合各维度奖励以强调优化不足的方面。实验结果表明,我们的方法(8B参数)在平衡准确率、SROCC和PLCC指标上达到了SOTA性能,优于Gemini-3.1、Qwen3.7-plus和GPT-5.5等强大的专有模型,同时可作为奖励引导虚拟试穿生成的有效优化信号。
英文摘要
Virtual try-on (VTON) requires not only realistic generation but also faithful preservation of garment characteristics. However, existing evaluation metrics such as PSNR, SSIM, KID and FID struggle to measure the consistency between the generated and reference garments, particularly in capturing the multi-dimensional characteristics of garment fidelity. To address this, we propose DAT: a Dimension-wise Assessment framework for virtual Try-on, which decomposes garment consistency into seven interpretable dimensions: silhouette, color, neckline and sleeve shape, major decoration and structure, material texture, fine-detail fidelity, and logo preservation, each formulated as a discrete attribute-level prediction task. To train this specialized assessment model, we adopt a two-stage learning paradigm comprising large-scale weak supervision on 50K samples, followed by refinement on 10K higher-quality annotations obtained via multi-model voting. Furthermore, we employ weighted cross-entropy loss to mitigate the severe label imbalance inherent across evaluation dimensions. Beyond its role as an evaluation framework, the assessment model can be integrated into reinforcement learning optimization of Qwen-Image-Edit for VTON, where dimension-wise rewards are adaptively aggregated to emphasize under-optimized aspects during training. Experimental results show that our method (8B parameters) achieves state-of-the-art performance in terms of balanced accuracy, SROCC, and PLCC, outperforming strong proprietary models such as Gemini-3.1, Qwen3.7-plus, and GPT-5.5, while also serving as an effective optimization signal for reward-guided VTON generation
发表机构
- Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团)
- Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。