AI 中文总结
InstructVVT是基于DiT的视频虚拟试穿框架,通过双层级参考调控方案无需推理时空间先验,在ViViD-S等数据集上优于现有开源方法,解决了现有方法依赖辅助先验的问题。
AI 中文摘要
视频虚拟试穿是一项约束性极强的编辑任务,需精准替换目标人物的服装,同时严格保留原视频的空间结构与时间动态。现有方法严重依赖手工制作的辅助空间先验(如掩码、姿态)进行编辑控制,但这些先验在无约束的真实场景视频中易失效,且常将丰富的视觉上下文压缩为不完整的结构信号。此外,标准重建目标无法完全捕捉试穿任务特有的人类偏好。为解决这些挑战,我们提出InstructVVT,这是一种基于扩散Transformer(DiT)的指令驱动、参考引导的视频虚拟试穿框架,推理时无需空间先验。我们的核心思路是通过双层级参考调控方案,直接从输入三元组(源视频、参考服装、指令)中恢复细粒度控制:具体而言,多模态大语言模型(MLLM)推断用于目标消歧与结构保留的语义编辑令牌,轻量级调控路径则显式注入服装的细粒度视觉细节。最后,我们设计了试穿专用奖励,并利用DiffusionNFT算法使模型与人类偏好对齐。在ViViD-S和TripVVT-Bench上的大量实验表明,尽管所需推理时控制更少,InstructVVT在服装保真度、结构保留和时间一致性方面仍优于最先进的开源方法。
英文摘要
Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial structure and temporal dynamics. Existing methods heavily rely on auxiliary handcrafted spatial priors (e.g., masks, poses) for editing control. However, these priors are prone to failure in unconstrained real-world videos and often compress rich visual context into incomplete structural signals. Furthermore, standard reconstruction objectives fail to fully capture try-on-specific human preferences. To address these challenges, we propose InstructVVT, an instruction-driven and reference-guided video virtual try-on framework based on a Diffusion Transformer (DiT) that operates without inference-time spatial priors. Our core insight is to recover fine-grained control directly from the input triplet (source video, reference garment, and instruction) via a dual-level reference conditioning scheme. Specifically, an MLLM infers semantic edit tokens for target disambiguation and structural preservation, while a lightweight conditioning pathway explicitly injects fine-grained visual garment details. Finally, we design a try-on-specific reward and utilize the DiffusionNFT algorithm to align the model with human preferences. Extensive experiments on ViViD-S and TripVVT-Bench demonstrate that InstructVVT outperforms state-of-the-art open-source methods in garment fidelity, structural preservation, and temporal consistency, despite requiring fewer inference-time controls.
Comments23 pages, 10 figures. Dingbao Shao and Song Wu contributed equally. Zili Yi is the corresponding author