arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TexTailor:通过自适应服装条件实现纹理保持的视频虚拟试穿

TexTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning

Zijing Qin, Jun Zhou, Ruicheng Zhang, Jiaqi Hou, Zunnan Xu, Ronghui Li, Zhenyu Xie, Xiu Li

arXiv 2609.39335首次发表:更新:

发表机构

Tsinghua University; Mohamed bin Zayed University of Artificial Intelligence(清华大学; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TexTailor提出基于视频扩散Transformer的框架,通过时间步自适应调制和帧对齐位置编码,解决高分辨率视频虚拟试穿中服装细节保留与时间一致性问题。

AI 中文摘要

视频虚拟试穿因其在数字时尚和智能电商中的广泛潜力而日益受到关注。然而,现有方法主要聚焦于低分辨率设置,在扩展到高分辨率场景时仍面临重大挑战。这些局限性可归因于两个主要因素:(1)对丰富服装参考信息的利用不足;(2)在跨模态交互过程中缺乏服装与视频表示之间的显式位置建模,这削弱了细粒度的局部对应关系。为解决这些问题,我们提出了TexTailor,一个基于预训练视频扩散Transformer的高保真视频虚拟试穿框架。具体而言,我们引入了一种时间步自适应调制机制,在去噪过程中动态调整服装视觉表示。我们进一步开发了一种帧对齐的位置编码策略,以加强服装与视频之间的对应关系,并采用多源注入设计以减少异构条件之间的干扰。在多个视频虚拟试穿基准(包括高分辨率Eevee数据集)上的大量实验表明,TexTailor在服装细节保留、时间一致性和整体视频质量方面取得了具有竞争力的性能。

英文摘要

Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing methods primarily focus on low-resolution settings and still face substantial challenges when extended to high-resolution scenarios. These limitations can be attributed to two main factors: (1) the insufficient utilization of rich garment reference information, and (2) the lack of explicit positional modeling between garment and video representations during cross-modal interaction, which weakens fine-grained local correspondence. To address these issues, we propose TexTailor, a high-fidelity video virtual try-on framework built upon a pretrained video Diffusion Transformer. Specifically, we introduce a timestep-adaptive modulation mechanism to dynamically adjust garment visual representations throughout denoising. We further develop a frame-aligned positional encoding strategy to strengthen garment-to-video correspondence, together with a multi-source injection design that reduces interference among heterogeneous conditions. Extensive experiments on multiple video virtual try-on benchmarks, including the high-resolution Eevee dataset, demonstrate that TexTailor achieves competitive performance in garment detail preservation, temporal consistency, and overall video quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑