发表机构
Tsinghua University; Zhongguancun Academy; South China University of Technology; Beijing Jiaotong University(清华大学; 中关村学院; 华南理工大学; 北京交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LiveVVT是一种滚动式流式扩散框架,通过保留双向建模与互补内存维持一致性,结合渐进式蒸馏优化,实现了比同等规模模型延迟降26倍、吞吐量提11倍的高保真实时视频虚拟试穿。
AI 中文摘要
基于扩散模型的视频虚拟试穿(VVT)通过双向时空建模实现了高视觉保真度,但完整片段依赖会在实际连续部署中导致过高的延迟和计算开销。单纯强制因果性会破坏预训练的双向先验,大幅降低合成质量。我们提出LiveVVT,一种滚动式流式扩散框架,可在因果循环生成中保留有限的双向建模能力。在固定大小窗口内,LiveVVT在有限前瞻范围内联合对多个视频块进行去噪,保留局部双向交互的同时每次迭代输出一个干净的视频块。窗口之外,两种互补内存维持长期一致性:有限的时间内存传播近期动态和遮挡上下文;持久的全局外观内存(从目标衣物和正面试穿关键帧一次性构建)在整个流中锚定衣物细节和试穿后外观。我们还提出渐进式蒸馏框架,整合双向VVT学习、用于因果少步适应的教师轨迹回归,以及协同匹配蒸馏(将教师分布匹配与真实视频上的滚动流匹配相结合,使优化与循环推理对齐)。在配对和非配对长序列基准上的实验表明,该方法比同等规模模型的生成质量更优,延迟降低26倍,吞吐量提高11倍,实现了高保真实时流式VVT。
英文摘要
Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. Within a fixed-size window, LiveVVT jointly denoises multiple video chunks under bounded look-ahead, preserving local bidirectional interactions while emitting one clean chunk per iteration. Beyond the window, two complementary memories sustain long-term consistency: a bounded temporal memory propagates recent dynamics and occlusion context, whereas a persistent global appearance memory, constructed once from the target garment and a frontal try-on keyframe, anchors garment details and dressed appearance throughout the stream. We further introduce a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference. Experiments on paired and unpaired long-sequence benchmarks demonstrate superior generation quality over similarly sized models, with $26\times$ lower latency and $11\times$ higher throughput, enabling high-fidelity real-time streaming VVT.
Comments16 pages, 13 figures,