AI 中文总结
UniVVT是一种统一端到端视频虚拟试穿框架,以多模态大语言模型为核心,采用三阶段渐进训练策略,在多基准上实现了优于主流方法的高保真试穿效果。
AI 中文摘要
视频虚拟试穿(Video Virtual Try-On,VVT)旨在合成人物穿着目标服装的视频,同时保留人物身份、动作和场景动态。主流方法将VVT视为掩码条件下的视频修复,依赖人体解析、姿态估计和服装扭曲等独立模块。这种多阶段设计增加了部署复杂度,更关键的是,显式几何先验的错误会不可逆地传播到生成的视频中。我们提出UniVVT,一种统一的端到端框架,将VVT重新定义为语义条件下的视频生成,推理阶段无需掩码、姿态和扭曲模块。其核心是基于多模态大语言模型构建的场景-任务感知器,将源视频、目标服装和任务指令联合编码为紧凑的、任务感知的潜在令牌,隐式捕捉需要转移的内容以及转移的位置和方式。随后,一个轻量级语义桥将这些令牌与扩散型视频生成器的条件空间对齐,实现连贯的服装转移。为了稳健耦合异构组件,我们设计了包含语义对齐、联合任务适配和灵活分辨率细化的三阶段渐进式训练策略。大量实验表明,UniVVT在多个基准上达到了最先进的性能,验证了隐式语义指导作为脆弱几何预处理的简单有效替代方案,适用于端到端虚拟试穿。
英文摘要
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into the generated video. We present UniVVT, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference. At its core, a scene-task perceiver built on a Multimodal Large Language Model jointly encodes the source video, target garment, and task instruction into compact, task-aware latent tokens, implicitly capturing what to transfer and where and how to transfer it. A lightweight semantic bridge then aligns these tokens with the conditioning space of a diffusion-based video generator, enabling coherent garment transfer. To robustly couple the heterogeneous components, we devise a three-stage progressive training strategy comprising semantic alignment, joint task adaptation, and flexible-resolution refinement. Extensive experiments demonstrate that UniVVT achieves state-of-the-art performance across multiple benchmarks, validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.
Comments17 pages,21 figures