发表机构
University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现实中服装分层搭配的虚拟试穿难题,提出LVTON方法。先通过自动数据生成管道训练获取一般VTON先验知识,再在专用数据集微调学习分层逻辑,在LVTON基准及传统VTON基准上取得优异成果,展现零样本能力。
AI 中文摘要
在现实世界中,时尚涉及分层搭配,如在衬衫外面加一件夹克,或进行一系列的添加和移除衣物层次的操作,而不仅仅是单层替换。这一现实任务在现有虚拟试穿(VTON)方法中仍是挑战,现有方法擅长单层替换,却非设计用于对现有服装进行分层或去层。本文提出分层虚拟试穿(LVTON),一种能保留现有服装同时实现顺序分层的基准和方法。研究发现当前VTON范式根本无法应对LVTON,因其依赖与布料无关的表示和单品数据集而丢弃了关键分层上下文。关键见解是将LVTON挑战分解为两个不同能力:一般VTON先验(如变形、身份保留)和特定分层知识(如分层顺序和遮挡推理)。首先通过在自动数据生成管道生成的数据上训练获取一般VTON先验知识,该管道通过分割和修复从时尚视频合成样本;其次在小型专用LVTON数据集上微调以学习分层逻辑。该方法在LVTON基准上取得了最优结果,并在传统VTON基准上展示了卓越的泛化能力,微调时创造了新的最优结果且展现出零样本能力。
英文摘要
In the real world, fashion is about layering: adding a jacket over a shirt, or a sequence of adding and removing layers, rather than just a single-layer swap. This fundamental real-world task remains a challenge in existing Virtual Try-On (VTON) methods, which excel at single-layer replacement but are not designed to layer or de-layer an existing outfit. This paper proposes Layering Virtual Try-On (LVTON), a layering benchmark and method that preserves an existing outfit while enabling sequential layering. We find that current VTON paradigms are fundamentally ill-equipped for LVTON, as their reliance on cloth-agnostic representations and single-item datasets discards essential layering context. Our key insight is that the LVTON challenge must be disentangled into two distinct competencies: (1) General VTON Priors (e.g., deformation, identity preservation) and (2) Specific Layering Knowledge (e.g., layering order and occlusion reasoning). First, our model obtains general VTON priors by being trained on data produced by an automatic data generation pipeline that synthesizes samples from fashion videos via segmentation and inpainting. Second, the model is fine-tuned on a small, dedicated LVTON dataset to learn the layering logic. Our method achieves state-of-the-art results on our LVTON benchmark and demonstrates superior generalizability on traditional VTON benchmarks, setting new state-of-the-art results when fine-tuned and exhibiting zero-shot capabilities.
CommentsAccepted to ECCV 2026