arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30450cs.CV

FlowVVTON:基于流引导的无掩码视频虚拟试穿

FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On

Shengyao Chen, Xianbing Sun, Liqing Zhang, Jianfu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

FlowVVTON是一种无掩码视频虚拟试穿框架,以光流为训练监督,通过两阶段训练实现多尺度时间一致性,在TikTokDress数据集上较基线方法(如SwiftTry)时间一致性提升显著,无需额外标注。

中文摘要 AI 辅助

视频虚拟试穿旨在将目标服装跨视频帧迁移到运动的人物身上。现有方法依赖人体解析掩码或姿态关键点,在大动作和遮挡场景下常失效,导致边界伪影和时间不一致。进一步的局限是,多数方法仅依赖注意力机制进行时间建模,缺乏显式运动监督。我们提出FlowVVTON,一种完全消除解析掩码依赖的无掩码框架。光流仅用作训练时的监督信号:在生成模型所有层应用光流扭曲潜在损失,通过在显式物理运动约束下对齐相邻帧特征,实现多尺度时间一致性。两阶段训练策略在引入流引导时间监督前建立无掩码空间对齐。在TikTokDress数据集上的实验显示,FlowVVTON大幅优于基线方法,尤其在时间一致性上较SwiftTry实现了5.7倍的VFID-R提升,且在任何阶段均无需分割掩码、姿态关键点或区域标注。

英文摘要

Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose keypoints that frequently fail under large motions and occlusions, causing boundary artifacts and temporal inconsistency. A further limitation is that most approaches rely solely on attention mechanisms for temporal modeling, providing no explicit motion supervision. We propose FlowVVTON, a mask-free framework that eliminates parsing mask dependency entirely. Optical flow is used solely as a training-time supervision signal: a flow-warped latent loss, applied across all layers of the generation model, enforces multi-scale temporal consistency by aligning adjacent-frame features under explicit physical motion constraints. A two-stage training strategy establishes mask-free spatial alignment before introducing flow-guided temporal supervision. Experiments on TikTokDress show that FlowVVTON outperforms baselines by substantial margins, particularly in temporal consistency (5.7$\times$ VFID-R improvement over SwiftTry), while requiring no segmentation masks, pose keypoints, or region annotations at any stage.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑