BooM-VVT:借助图像级伪数据提升无掩码视频虚拟试穿性能
BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data
浏览论文内容
中文总结 AI 辅助
针对现有无掩码视频虚拟试穿依赖掩码、成本高及服装一致性差的问题,提出BooM-VVT框架,采用图像级伪数据、服装敏感关键帧采样等技术,构建OmniView数据集,实现更优的时间一致性与服装保真度。
中文摘要 AI 辅助
视频虚拟试穿(Video Virtual Try-On,VVT)旨在生成人物穿着目标服装的逼真视频。近期方法利用关键帧驱动的视频生成范式提升野外场景性能,但仍依赖掩码定位试穿区域,易受大幅动作和严重遮挡影响。尽管基于图像的无掩码试穿方法通过大规模伪数据取得良好效果,但将该范式扩展至视频仍存在困难,因为构建视频级伪数据成本过高。此外,粗糙的关键帧采样和多视角试穿数据的稀缺性,限制了现有关键帧驱动方法在保持服装一致性和处理多样试穿任务方面的能力。为解决这些挑战,我们提出BooM-VVT,一种基于关键帧驱动范式的无掩码VVT框架。为实现无掩码VVT,我们引入多阶段训练策略,利用图像级伪数据进行无掩码定位学习,大幅减少对成本高昂的视频级伪数据的需求。为提升服装一致性,我们提出服装敏感关键帧采样方法,基于与服装相关的身体区域选择关键帧,以更好地捕捉服装外观。我们进一步引入帧共享3D-RoPE,建立关键帧与目标视频帧之间的时空对应关系,实现准确的服装细节迁移。最后,我们构建OmniView,一个大规模多视角试穿数据集,以支持复杂相机视角和多样试穿任务下的可靠试穿视频生成。大量实验表明,BooM-VVT相较于现有方法,在时间一致性和服装保真度方面表现更优。项目页面:this https URL
英文摘要
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent methods leverage a keyframe-driven video generation paradigm to improve in-the-wild performance, yet they still rely on masks to localize try-on regions, making them vulnerable to large motions and severe occlusions. Although mask-free image-based try-on methods have shown promising results by leveraging large-scale pseudo data, extending this paradigm to videos remains difficult, as constructing video-level pseudo data is prohibitively expensive. Furthermore, coarse keyframe sampling and the scarcity of multi-view try-on data limit existing keyframe-driven methods in maintaining garment consistency and handling diverse try-on tasks. To address these challenges, we propose BooM-VVT, a mask-free VVT framework built upon the keyframe-driven paradigm. To achieve mask-free VVT, we introduce a multi-stage training strategy that leverages image-level pseudo data for mask-free localization learning, substantially reducing the need for costly video-level pseudo data. To improve garment consistency, we propose Garment-Sensitive Keyframe Sampling, which selects keyframes based on garment-relevant body regions to better capture garment appearance. We further introduce Frame-Shared 3D-RoPE to establish spatiotemporal correspondences between keyframes and target video frames for accurate garment-detail transfer. Finally, we construct OmniView, a large-scale multi-view try-on dataset to support reliable try-on video generation under complex camera viewpoints and diverse try-on tasks. Extensive experiments demonstrate that BooM-VVT achieves superior temporal consistency and garment fidelity over existing methods. Project page: https://boomvvt.github.io/boomvvt.
发表机构
- Nanjing University of Science and Technology(南京理工大学)
- University of Science and Technology of China(中国科学技术大学)
- National University of Singapore(新加坡国立大学)
- Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。